SGWAFusion: A semantic-guided wavelet attention fusion network for infrared and visible image fusion
Infrared and visible image fusion aims to generate a more informative fused image by integrating the salient target responses of infrared images with the rich texture and structural information of visible images. Although existing deep learning-based methods have achieved promising results, most of them still mainly rely on spatial-domain fusion and low-level visual constraints, making it difficult to simultaneously preserve frequency-domain details, target saliency, and high-level semantic consistency in complex scenes. To address this issue, we propose SGWAFusion, a Semantic-Guided Wavelet Attention Fusion Network. The proposed method adopts modality-specific hierarchical encoders to separately extract visible texture-structure features and infrared salient target features, and introduces discrete wavelet transform and a Multi-Scale Gated Convolutional Representation Module to enhance the frequency-domain structural modeling capability of deep features. In the deep fusion stage, the Semantic-Guided Attention Fusion Module employs a frozen SigLIP vision encoder to provide scene-level semantic conditions, which are converted by FiLM into channel-wise modulation and further processed by a Spatial-Channel Attention Block to refine the spatial and channel responses of the joint features, thereby promoting semantic consistency while preserving infrared saliency and visible structural details. Finally, the fused deep features are integrated with shallow aggregated details and fed into a decoder to reconstruct the final fused image. Extensive experiments on the TNO, M 3 FD, and RoadScene datasets demonstrate that SGWAFusion achieves competitive performance in terms of visual quality, six objective metrics, and downstream tasks, while ablation studies further verify the effectiveness of wavelet-domain modeling, semantic modulation, and spatial-channel attention refinement.
Authors
- Shihao Liu (ORCID: https://orcid.org/0000-0002-0645-5319)
- Jiawei Liu (ORCID: https://orcid.org/0000-0002-6105-8121)
- Guiling Sun (ORCID: https://orcid.org/0000-0001-5283-1760)
- Wenrui Zhang (ORCID: https://orcid.org/0009-0008-4644-3356)
Institutions
- Nankai University (CN)
Publication Details
- Journal
- Optics & Laser Technology
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1016/j.optlastec.2026.116488
- Primary Topic
- Advanced Image Fusion Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00