MS-GFFE: Multi-stage global feature fusion and enhancement method for lesion segmentation
Lesion segmentation in medical images remains challenging because of substantial scale variation, ambiguous boundaries, and high visual similarity between foreground lesions and surrounding tissues. Existing methods often struggle to preserve fine-grained local details while effectively modeling global contextual dependencies. To address these challenges, we propose a Multi-stage Global Feature Fusion and Enhancement Network (MS-GFFE) for medical image lesion segmentation. MS-GFFE adopts a U-shaped CNN–Transformer hybrid architecture and introduces a stage-specific collaboration mechanism that integrates local feature extraction, global dependency modeling, cross-scale semantic alignment, and spatial-detail reconstruction. Specifically, the Multi-scale Dilated Convolution Residual block (MDR) employs parallel dilated convolutions to construct multi-receptive-field representations, thereby improving the characterization of lesions at different scales. The Selective Dense Connection block (SDC) adaptively filters densely reused Swin features to suppress redundant responses and retain informative global representations. The Multi-scale Information Cross-fusion block (MIC) aligns encoder and decoder features through cross-branch channel interaction, reducing the semantic gap in skip connections. The Detail Enhancement Upsampling block (DEU) enhances boundary and small-lesion information before spatial-resolution recovery, thereby mitigating detail degradation caused by repeated interpolation. MS-GFFE was systematically evaluated on five public datasets, namely ISIC2018, PH2, DRIVE, STARE, and CHASE_DB1, as well as an internal multimodal endoscopic dataset comprising 3,960 images. On the internal dataset, MS-GFFE-L achieved a mean Intersection over Union (mIoU) of 0.6063, a Dice Similarity Coefficient (DSC) of 0.7355, and a specificity of 0.9371. Controlled ablation experiments showed that the complete configuration increased mIoU from 0.5296 to 0.5742, corresponding to a relative improvement of 8.42% over the baseline and confirming the effectiveness of the stage-specific collaborative design. The MS-GFFE-S, MS-GFFE-M, and MS-GFFE-L variants contained 100.02 M, 286.95 M, and 400.08 M parameters and required 178.36, 512.61, and 589.65 GFLOPs, respectively, explicitly revealing the trade-off between segmentation accuracy and computational complexity across model capacities. These results indicate that MS-GFFE effectively coordinates local structural details with global contextual information across the evaluated imaging modalities and segmentation tasks, providing a competitive framework for segmenting complex lesions and fine anatomical structures.
Authors
- Jiamin Qin (ORCID: https://orcid.org/0000-0002-3692-6353)
- Zhenxiang He
- Xia Zhou
- Yuling Chen
- Qiang Liu
Institutions
- Southwestern University of Finance and Economics (CN)
- Mianyang Normal University (CN)
- Sichuan Mianyang 404 Hospital (CN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-04
- DOI
- https://doi.org/10.1038/s41598-026-74003-7
- Primary Topic
- Medical Image Segmentation Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00