Multi-scale Coarse-to-fine Alignment Network for weakly aligned Red–Green–Blue and Thermal salient object detection
Red–Green–Blue and Thermal (RGBT) salient object detection (SOD) aims to detect common conspicuous regions from the visible light and thermal infrared image pair. Due to the different imaging mechanisms of visible light and thermal images, they are usually weakly aligned. Most existing works need to align two modalities manually, which leads to large manual cost, while alignment-based methods can relieve the cost but are difficult to achieve accurate alignment between two modalities due to a large modality gap. To address these problems, we propose a Multi-scale Coarse-to-fine Alignment Network (MCANet), which pursues the accurate cross-modal alignment by incorporating the multi-scale information and the coarse-to-fine scheme, for weakly aligned RGBT SOD. In particular, we design a coarse-to-fine alignment module (CAM) based on the affine transform networks and deformable convolutions to achieve accurate cross-modal alignment in a progressive manner. To leverage the rich features available at different levels, we incorporate the CAM into each stage of the detection network, and employs a multi-scale consistency loss to constrain the affine transformations predicted at different feature levels, thereby reducing inconsistent geometric estimates and promoting stable optimization.In addition, we design a cross-modal cross-level fusion module (CCFM) that utilizes the attention mechanism to fuse the features before and after alignment in a top-down strategy. We conduct extensive experiments on three public benchmark datasets, and the results show that our method achieves comparable performance compared to state-of-the-art methods.
Authors
- Zhengzheng Tu (ORCID: https://orcid.org/0000-0002-9689-8657)
- Lili Huang (ORCID: https://orcid.org/0000-0002-7085-7646)
- Feifan Sun
- Danying Lin
- Chenglong Li
Institutions
- Anhui University (CN)
Publication Details
- Journal
- Engineering Applications of Artificial Intelligence
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1016/j.engappai.2026.116394
- Primary Topic
- Visual Attention and Saliency Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00