LMCAN: A Lightweight Multiscale Contextual Attention Network for Hyperspectral and Multispectral Image Fusion
Hyperspectral and multispectral image fusion requires enhancing spatial details while preserving the pixel-wise spectral fidelity of hyperspectral observations. Transformer-based approaches provide effective contextual modeling, yet hierarchical token merging may cause spectral mixing, whereas pixel-preserving tokenization restricts the spatial context captured by window attention. To resolve this trade-off, we propose a Lightweight Multiscale Contextual Attention Network (LMCAN) that reconstructs spatial context without altering the pixel-level representation. The proposed network progressively broadens contextual perception and coordinates complementary spatial dependencies within the attention process, enabling information at different scales to interact adaptively rather than being modeled in isolation. It further recovers local correlations omitted by fixed window partitioning, improving the reconstruction of boundaries and fine structures. Through this unified design, spectral preservation and spatial context restoration are jointly achieved within a shallow and efficient architecture. Experiments on four benchmark datasets demonstrate competitive accuracy with substantially lower complexity. On CAVE, LMCAN achieves 49.04 dB PSNR and 2.47 SAM with only 0.157 M parameters and 12.06 G FLOPs.
Authors
- Tianyi Xu (ORCID: https://orcid.org/0000-0001-9793-4894)
- Mingming Ma (ORCID: https://orcid.org/0000-0002-1365-3559)
- Yi Niu
- Jialiang Wu
- Peixian He
- Zhengzhong Fang
Institutions
- Xidian University (CN)
- Peng Cheng Laboratory (CN)
Publication Details
- Journal
- Remote Sensing
- Published
- 2026-09-14
- DOI
- https://doi.org/10.3390/rs18183154
- Primary Topic
- Advanced Image Fusion Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Natural Science Foundation of China
- China Postdoctoral Science Foundation