MA-YOLO for robust multimodal feature fusion under weak cross-modal alignment in remote sensing object detection

Abstract Multimodal remote sensing object detection benefits from the complementary characteristics of visible and infrared modalities. However, owing to the heterogeneous imaging mechanisms of RGB and infrared sensors, feature representations from the two modalities progressively become semantically inconsistent during hierarchical feature extraction, even when the input image pairs are geometrically registered. Such feature-level semantic inconsistency weakens the robustness of multimodal feature fusion and ultimately degrades object localization accuracy. To address this issue, we propose Misalignment-Aware YOLO (MA-YOLO). MA-YOLO incorporates a Global Semantic Encoder (GSE) to capture long-range contextual dependencies, a Robust Adaptive Fusion (RAF) module to adaptively refine multiscale feature representations by suppressing spatially inconsistent responses during hierarchical feature propagation, and an NWD-assisted bounding-box regression strategy to improve localization robustness under spatial uncertainty and ambiguous object boundaries. Extensive experiments on the VEDAI benchmark demonstrate that MA-YOLO achieves an average mAP $$_{50}$$ of 76.93% under tenfold cross-validation. Spatial perturbation experiments further show that the proposed framework maintains relatively stable detection performance under cross-modal shifts of up to 32 pixels, supporting its robustness to cross-modal spatial inconsistency. Additional experiments on the DIOR and NWPU VHR-10 datasets also demonstrate the applicability of the proposed framework to single-modal optical remote sensing scenarios.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-28
DOI
https://doi.org/10.1038/s41598-026-72060-6
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

MA-YOLO for robust multimodal feature fusion under weak cross-modal alignment in remote sensing object detection

Yajun Xie, Jianqiong Huang, Fang Lu, Zhihong Lin et al.
Scientific Reports
Advanced Neural Network Applications
article

MA-YOLO for robust multimodal feature fusion under weak cross-modal alignment in remote sensing object detection

Yajun Xie, Jianqiong Huang, Fang Lu, Zhihong Lin, Dandan Yao, Xiaolu Meng, Waqar Khan
article en

Abstract

Abstract Multimodal remote sensing object detection benefits from the complementary characteristics of visible and infrared modalities. However, owing to the heterogeneous imaging mechanisms of RGB and infrared sensors, feature representations from the two modalities progressively become semantically inconsistent during hierarchical feature extraction, even when the input image pairs are geometrically registered. Such feature-level semantic inconsistency weakens the robustness of multimodal feature fusion and ultimately degrades object localization accuracy. To address this issue, we propose Misalignment-Aware YOLO (MA-YOLO). MA-YOLO incorporates a Global Semantic Encoder (GSE) to capture long-range contextual dependencies, a Robust Adaptive Fusion (RAF) module to adaptively refine multiscale feature representations by suppressing spatially inconsistent responses during hierarchical feature propagation, and an NWD-assisted bounding-box regression strategy to improve localization robustness under spatial uncertainty and ambiguous object boundaries. Extensive experiments on the VEDAI benchmark demonstrate that MA-YOLO achieves an average mAP $$_{50}$$ of 76.93% under tenfold cross-validation. Spatial perturbation experiments further show that the proposed framework maintains relatively stable detection performance under cross-modal shifts of up to 32 pixels, supporting its robustness to cross-modal spatial inconsistency. Additional experiments on the DIOR and NWPU VHR-10 datasets also demonstrate the applicability of the proposed framework to single-modal optical remote sensing scenarios.

Scientific Reports
Fuzhou University of International Studies and Trade (CN), Fuzhou University (CN)
Openalex Percentile: Top 14%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

MA-YOLO for robust multimodal feature fusion under weak cross-modal alignment in remote sensing object detection — Yajun Xie, Jianqiong Huang, et al. · Scientific Reports (2026) | TGRS Research Map | TGRS