MSFEMamba: mamba-based multi-scale feature fusion and foreground enhancement for semantic segmentation of remote sensing images
Semantic segmentation of remote sensing images is vital for urban planning and environmental monitoring. However, despite progress with CNNs and Transformers, complex land cover and boundary interference in high-resolution images still challenge the spatial representation and long-range dependency modelling of existing models. To address these issues, we propose MSFEMamba, a novel Mamba-based semantic segmentation network integrating multi-scale feature fusion and foreground enhancement. In the encoder, a dual‑branch context aggregation (DBCA) module is designed to efficiently model multi-scale global context. DBCA introduces a self-attention mechanism and a state space model (SSM) in parallel to capture long-range dependencies in both spatial and sequential dimensions. Within DBCA, a channel–spatial fusion module (CSFM) adaptively fuses spatial context from LSRFormer and sequential context from Mamba via channel and spatial attention mechanisms. In the decoder, a foreground enhancement module (FEM) selectively strengthens cross-level features through attention gating. This enhances foreground regions and boundary responses while preserving low-level details, thereby improving the segmentation of small targets and complex boundaries. Experiments on the Vaihingen, Potsdam and LoveDA datasets demonstrate that MSFEMamba achieves mIoU scores of 84.90%, 87.93% and 55.23%, respectively.
Authors
- Zhong Xingyu
- Ying Xia
- Jiangfan Feng
Institutions
- Chongqing University of Posts and Telecommunications (CN)
- Yangtze Normal University (CN)
Publication Details
- Journal
- International Journal of Image and Data Fusion
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1080/19479832.2026.2738456
- Primary Topic
- Remote-Sensing Image Classification
- Type
- article
- Field-Weighted Citation Impact
- 0.00