Automated lung tumor segmentation using enhanced U-Net architectures: a multi-backbone comparative study for clinical applications
Lung cancer remains a leading cause of cancer-related mortality worldwide, and manual delineation of tumors on computed tomography (CT) is time-consuming and prone to inter-observer variability. This study systematically evaluates nine deep learning configurations for automated lung tumor segmentation across the full spectrum of tumor sizes, providing evidence-based guidance for architectural selection in clinical applications. Three encoder-decoder architectures (U-Net, U-Net++, UNet 3+) were combined with three pre-trained backbones (VGG16, ResNet50, Xception), yielding nine configurations. A combined dataset of 161 subjects (63 public Kaggle cases, 1,647 slices; 98 patients from Imam Reza Hospital, Tabriz, 500 slices, collected retrospectively in 2025) was preprocessed, yielding 1,322 tumor-containing CT slices standardized to 128 × 128 pixels. A morphological pipeline (CLAHE, Otsu thresholding, automated lung extraction) isolated pulmonary structures. Models were trained with combined Dice-Focal loss and Adam optimizer via 5-fold stratified cross-validation (80/20 split). Tumors were also retrospectively stratified into five size categories (diameter approximately 0.6–30.2 mm) and evaluated using both generalist and independently trained size-specific models. Seven metrics (IoU, Dice, sensitivity, specificity, precision, accuracy, F1-score) were computed, with significance evaluated via ANOVA and Kruskal-Wallis tests. U-Net with VGG16 achieved the highest overall performance (IoU: 0.8838 ± 0.0229, Dice: 0.9342 ± 0.0164), while U-Net + + with Xception and UNet 3 + with VGG16 showed superior stability (Dice SD: 0.0009). UNet 3 + with ResNet50 attained the highest sensitivity (0.9338 ± 0.0142) and precision (0.9449 ± 0.0059). All configurations achieved specificity ≥ 0.9991 and accuracy ≥ 0.9980 (ANOVA p < 0.001; Kruskal-Wallis p < 0.001 for IoU/Accuracy, p < 0.01 for Dice). Performance improved with tumor size, from IoU 0.7912 ± 0.0287 (smallest category) to 0.9106 ± 0.0131 (largest), with size-specific training yielding only marginal gains (0.3–1.3% points) over the generalist model. The systematic architectural comparison, including size-stratified evaluation, provides evidence-based guidance for model selection in lung tumor segmentation, with morphological preprocessing effectively reducing false positives, supporting applications in radiation therapy planning and computer-aided detection.
Authors
- Hossein Najaf-Zadeh (ORCID: https://orcid.org/0000-0001-7411-8480)
- Elaheh Asghari
- Peren Jerfi Canatalay (ORCID: https://orcid.org/0000-0002-0702-2179)
- Rabehe ostadhassan
- Alireza Asghari (ORCID: https://orcid.org/0009-0000-2905-2320)
Institutions
- Tabriz University of Medical Sciences (IR)
- Politecnico di Torino (IT)
- Wichita State University (US)
- Islamic Azad University Central Tehran Branch (IR)
- Istanbul University (TR)
Publication Details
- Journal
- BioData Mining
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1186/s13040-026-00604-7
- Primary Topic
- Lung Cancer Diagnosis and Treatment
- Type
- article
- Field-Weighted Citation Impact
- 0.00