SGWAFusion: A semantic-guided wavelet attention fusion network for infrared and visible image fusion

Infrared and visible image fusion aims to generate a more informative fused image by integrating the salient target responses of infrared images with the rich texture and structural information of visible images. Although existing deep learning-based methods have achieved promising results, most of them still mainly rely on spatial-domain fusion and low-level visual constraints, making it difficult to simultaneously preserve frequency-domain details, target saliency, and high-level semantic consistency in complex scenes. To address this issue, we propose SGWAFusion, a Semantic-Guided Wavelet Attention Fusion Network. The proposed method adopts modality-specific hierarchical encoders to separately extract visible texture-structure features and infrared salient target features, and introduces discrete wavelet transform and a Multi-Scale Gated Convolutional Representation Module to enhance the frequency-domain structural modeling capability of deep features. In the deep fusion stage, the Semantic-Guided Attention Fusion Module employs a frozen SigLIP vision encoder to provide scene-level semantic conditions, which are converted by FiLM into channel-wise modulation and further processed by a Spatial-Channel Attention Block to refine the spatial and channel responses of the joint features, thereby promoting semantic consistency while preserving infrared saliency and visible structural details. Finally, the fused deep features are integrated with shallow aggregated details and fed into a decoder to reconstruct the final fused image. Extensive experiments on the TNO, M 3 FD, and RoadScene datasets demonstrate that SGWAFusion achieves competitive performance in terms of visual quality, six objective metrics, and downstream tasks, while ablation studies further verify the effectiveness of wavelet-domain modeling, semantic modulation, and spatial-channel attention refinement.

Authors

Institutions

Publication Details

Journal
Optics & Laser Technology
Published
2026-09-28
DOI
https://doi.org/10.1016/j.optlastec.2026.116488
Primary Topic
Advanced Image Fusion Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SGWAFusion: A semantic-guided wavelet attention fusion network for infrared and visible image fusion

Shihao Liu, Jiawei Liu, Guiling Sun, Wenrui Zhang
Optics & Laser Technology
Advanced Image Fusion Techniques
article

SGWAFusion: A semantic-guided wavelet attention fusion network for infrared and visible image fusion

Shihao Liu, Jiawei Liu, Guiling Sun, Wenrui Zhang
article en

Abstract

Infrared and visible image fusion aims to generate a more informative fused image by integrating the salient target responses of infrared images with the rich texture and structural information of visible images. Although existing deep learning-based methods have achieved promising results, most of them still mainly rely on spatial-domain fusion and low-level visual constraints, making it difficult to simultaneously preserve frequency-domain details, target saliency, and high-level semantic consistency in complex scenes. To address this issue, we propose SGWAFusion, a Semantic-Guided Wavelet Attention Fusion Network. The proposed method adopts modality-specific hierarchical encoders to separately extract visible texture-structure features and infrared salient target features, and introduces discrete wavelet transform and a Multi-Scale Gated Convolutional Representation Module to enhance the frequency-domain structural modeling capability of deep features. In the deep fusion stage, the Semantic-Guided Attention Fusion Module employs a frozen SigLIP vision encoder to provide scene-level semantic conditions, which are converted by FiLM into channel-wise modulation and further processed by a Spatial-Channel Attention Block to refine the spatial and channel responses of the joint features, thereby promoting semantic consistency while preserving infrared saliency and visible structural details. Finally, the fused deep features are integrated with shallow aggregated details and fed into a decoder to reconstruct the final fused image. Extensive experiments on the TNO, M 3 FD, and RoadScene datasets demonstrate that SGWAFusion achieves competitive performance in terms of visual quality, six objective metrics, and downstream tasks, while ablation studies further verify the effectiveness of wavelet-domain modeling, semantic modulation, and spatial-channel attention refinement.

Optics & Laser TechnologyVol. 204
Nankai University (CN)
Openalex Percentile: Top 14%
Advanced Image Fusion Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.