Vital-Net: Vision Integrated Transformer and Attention Network for Lung Nodule Segmentation on Full-Scale Images
摘要
Lung cancer, a leading cause of global cancer mortality, demands precise and early diagnostic methods to improve patient outcomes. Lung nodule segmentation plays a critical role in early detection, yet current models primarily trained on nodule-centered patches struggle with whole-image generalization. This study proposes a preprocessing pipeline, together with an enhanced residual U-Net model incorporating Atrous Spatial Pyramid Pooling (ASPP), spatial and channel attention (scSE), and a Vision Transformer (ViT) bottleneck to capture both local and global dependencies. To be specific, our approach addresses the limitations of patch-based methods by introducing a preprocessing pipeline that maintains entire image context, filtering nodules by malignancy and size to ensure clinical relevance. The model was trained and validated on nodule-containing data, with performance evaluated on both nodule-containing and clean (non-nodule) datasets. The top configurations, utilizing scSE in encoder layers and ViT in the bottleneck, achieved superior generalization and accuracy. This flexible architecture offers a robust solution for whole-image lung nodule segmentation, enhancing adaptability to unseen data and holding promise for real-world clinical applications.