XoFTR++: dual-stage cross-modal enhancement for visible-thermal point correspondence matching

Visible-light-to-thermal-infrared feature matching is essential for multi-sensor fusion but remains challenging because of substantial modality discrepancies in image-formation mechanisms, texture statistics, and structural responses. Existing coarse-to-fine frameworks such as XoFTR improve cross-modal matching, yet their cross-modal interaction is concentrated primarily in the Transformer matching stage, leaving stage-wise error propagation insufficiently addressed. This study presents XoFTR++, a dual-stage cross-modal enhancement framework for visible–thermal correspondence matching. At the coarse stage, Cross-Modal Feature Enhancement (CMFE), Adaptive Feature Pyramid Fusion (AFPF), and Coarse-Level Transformer Enhancement (CLTE) are introduced to reduce early modality shifts, inject multi-scale semantic context into fine features, and improve coarse-candidate reliability. At the fine stage, Fine-Level Cross-Modal Attention Enhancement (FCAE) models local context, modality differences, and bidirectional window interactions to reduce local matching ambiguity before final similarity estimation. On METU-VisTIR under RANSAC = 1.5, XoFTR + + improves AUC@5/10/20 from 12.31 ± 0.62/28.03 ± 0.52/44.92 ± 0.49 to 15.07 ± 0.44/29.95 ± 1.32/47.14 ± 1.28 in Sunny–Cloudy scenes, and from 22.05 ± 1.72/39.11 ± 1.63/54.96 ± 1.59 to 23.28 ± 0.52/41.02 ± 1.04/57.13 ± 1.21 in Cloudy–Cloudy scenes, with a 10.29% parameter increase. Controlled ablations and stage-wise diagnostics support the complementary contributions of the four modules. A separate synthetic-homography evaluation on the public FusionDN RoadScene and MSRS datasets improves aggregate AUC@5/10/20 and reduces mean corner error, providing supplementary transfer evidence under the specified external protocol.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-08-31
DOI
https://doi.org/10.1038/s41598-026-68975-9
Primary Topic
Advanced Image and Video Retrieval Techniques
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

XoFTR++: dual-stage cross-modal enhancement for visible-thermal point correspondence matching

Xinghui Zhu, Zichun Peng, Tianmin Li, Bo Qiao et al.
Scientific Reports
Advanced Image and Video Retrieval Techniques
article

XoFTR++: dual-stage cross-modal enhancement for visible-thermal point correspondence matching

Xinghui Zhu, Zichun Peng, Tianmin Li, Bo Qiao, Hongyan Zhang, Yidi Huang, Yiming Chen
article en

Abstract

Visible-light-to-thermal-infrared feature matching is essential for multi-sensor fusion but remains challenging because of substantial modality discrepancies in image-formation mechanisms, texture statistics, and structural responses. Existing coarse-to-fine frameworks such as XoFTR improve cross-modal matching, yet their cross-modal interaction is concentrated primarily in the Transformer matching stage, leaving stage-wise error propagation insufficiently addressed. This study presents XoFTR++, a dual-stage cross-modal enhancement framework for visible–thermal correspondence matching. At the coarse stage, Cross-Modal Feature Enhancement (CMFE), Adaptive Feature Pyramid Fusion (AFPF), and Coarse-Level Transformer Enhancement (CLTE) are introduced to reduce early modality shifts, inject multi-scale semantic context into fine features, and improve coarse-candidate reliability. At the fine stage, Fine-Level Cross-Modal Attention Enhancement (FCAE) models local context, modality differences, and bidirectional window interactions to reduce local matching ambiguity before final similarity estimation. On METU-VisTIR under RANSAC = 1.5, XoFTR + + improves AUC@5/10/20 from 12.31 ± 0.62/28.03 ± 0.52/44.92 ± 0.49 to 15.07 ± 0.44/29.95 ± 1.32/47.14 ± 1.28 in Sunny–Cloudy scenes, and from 22.05 ± 1.72/39.11 ± 1.63/54.96 ± 1.59 to 23.28 ± 0.52/41.02 ± 1.04/57.13 ± 1.21 in Cloudy–Cloudy scenes, with a 10.29% parameter increase. Controlled ablations and stage-wise diagnostics support the complementary contributions of the four modules. A separate synthetic-homography evaluation on the public FusionDN RoadScene and MSRS datasets improves aggregate AUC@5/10/20 and reduces mean corner error, providing supplementary transfer evidence under the specified external protocol.

Scientific Reports
Hunan Agricultural University (CN)
Hunan Agricultural University
Openalex Percentile: Top 13%
Advanced Image and Video Retrieval Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

XoFTR++: dual-stage cross-modal enhancement for visible-thermal point correspondence matching — Xinghui Zhu, Zichun Peng, et al. · Scientific Reports (2026) | TGRS Research Map | TGRS