Multi-scale Coarse-to-fine Alignment Network for weakly aligned Red–Green–Blue and Thermal salient object detection

Red–Green–Blue and Thermal (RGBT) salient object detection (SOD) aims to detect common conspicuous regions from the visible light and thermal infrared image pair. Due to the different imaging mechanisms of visible light and thermal images, they are usually weakly aligned. Most existing works need to align two modalities manually, which leads to large manual cost, while alignment-based methods can relieve the cost but are difficult to achieve accurate alignment between two modalities due to a large modality gap. To address these problems, we propose a Multi-scale Coarse-to-fine Alignment Network (MCANet), which pursues the accurate cross-modal alignment by incorporating the multi-scale information and the coarse-to-fine scheme, for weakly aligned RGBT SOD. In particular, we design a coarse-to-fine alignment module (CAM) based on the affine transform networks and deformable convolutions to achieve accurate cross-modal alignment in a progressive manner. To leverage the rich features available at different levels, we incorporate the CAM into each stage of the detection network, and employs a multi-scale consistency loss to constrain the affine transformations predicted at different feature levels, thereby reducing inconsistent geometric estimates and promoting stable optimization.In addition, we design a cross-modal cross-level fusion module (CCFM) that utilizes the attention mechanism to fuse the features before and after alignment in a top-down strategy. We conduct extensive experiments on three public benchmark datasets, and the results show that our method achieves comparable performance compared to state-of-the-art methods.

Authors

Institutions

Publication Details

Journal
Engineering Applications of Artificial Intelligence
Published
2026-10-07
DOI
https://doi.org/10.1016/j.engappai.2026.116394
Primary Topic
Visual Attention and Saliency Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Multi-scale Coarse-to-fine Alignment Network for weakly aligned Red–Green–Blue and Thermal salient object detection

Zhengzheng Tu, Lili Huang, Feifan Sun, Danying Lin et al.
Engineering Applications of Artificial Intelligence
Visual Attention and Saliency Detection
article

Multi-scale Coarse-to-fine Alignment Network for weakly aligned Red–Green–Blue and Thermal salient object detection

Zhengzheng Tu, Lili Huang, Feifan Sun, Danying Lin, Chenglong Li
article en

Abstract

Red–Green–Blue and Thermal (RGBT) salient object detection (SOD) aims to detect common conspicuous regions from the visible light and thermal infrared image pair. Due to the different imaging mechanisms of visible light and thermal images, they are usually weakly aligned. Most existing works need to align two modalities manually, which leads to large manual cost, while alignment-based methods can relieve the cost but are difficult to achieve accurate alignment between two modalities due to a large modality gap. To address these problems, we propose a Multi-scale Coarse-to-fine Alignment Network (MCANet), which pursues the accurate cross-modal alignment by incorporating the multi-scale information and the coarse-to-fine scheme, for weakly aligned RGBT SOD. In particular, we design a coarse-to-fine alignment module (CAM) based on the affine transform networks and deformable convolutions to achieve accurate cross-modal alignment in a progressive manner. To leverage the rich features available at different levels, we incorporate the CAM into each stage of the detection network, and employs a multi-scale consistency loss to constrain the affine transformations predicted at different feature levels, thereby reducing inconsistent geometric estimates and promoting stable optimization.In addition, we design a cross-modal cross-level fusion module (CCFM) that utilizes the attention mechanism to fuse the features before and after alignment in a top-down strategy. We conduct extensive experiments on three public benchmark datasets, and the results show that our method achieves comparable performance compared to state-of-the-art methods.

Engineering Applications of Artificial IntelligenceVol. 185
Anhui University (CN)
Openalex Percentile: Top 15%
Visual Attention and Saliency Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Multi-scale Coarse-to-fine Alignment Network for weakly aligned Red–Green–Blue and Thermal salient object detection — Zhengzheng Tu, Lili Huang, et al. · Engineering Applications of Artificial Intelligence (2026) | TGRS Research Map | TGRS