Boundary-Guided Dual-Perspective Cross-Modal Fusion Network for RGB-IR Object Detection

Visible-infrared (RGB-IR) object detection leverages multimodal information to ensure reliable perception in complex environments. However, dynamic scenes pose significant challenges due to the frequent inconsistency between scene-level modality contributions and local spatial reliability. Furthermore, standard feature extraction progressively attenuates boundary-sensitive structural cues, and unified fusion strategies often fail to capture spatially varying cross-modal complementarity. To overcome these limitations, we propose a Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives. Specifically, a Geometric Boundary Enhancement Module (GBEM) embeds Sobel-based high-frequency priors into shallow dual-modal features via residual spatial modulation, preventing the loss of crucial localization cues during downsampling. In the deep semantic space, a Hybrid Dual-Perspective Adaptive Fusion Module (HDAM) employs an illumination-aware branch for global modality weighting and a spatial confidence-driven branch for local cross-modal rectification. A spatial gating mechanism then dynamically reconciles these macro-environmental and micro-signal features. Extensive experiments on M3FD, LLVIP, and DroneVehicle demonstrate the effectiveness of BDPNet. Compared with state-of-the-art methods, BDPNet improves mAP50-95 by 0.8% and 1.0% on M3FD and LLVIP, respectively, and improves mAP50 by 0.6% on DroneVehicle, while using substantially fewer parameters and lower computational cost.

Authors

Institutions

Publication Details

Journal
Remote Sensing
Published
2026-09-15
DOI
https://doi.org/10.3390/rs18183175
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Boundary-Guided Dual-Perspective Cross-Modal Fusion Network for RGB-IR Object Detection

Guirong Feng, Xiumei Chen, Zhiwei Fu, Huachen Lin
Remote Sensing
Advanced Neural Network Applications
article

Boundary-Guided Dual-Perspective Cross-Modal Fusion Network for RGB-IR Object Detection

Guirong Feng, Xiumei Chen, Zhiwei Fu, Huachen Lin
article en

Abstract

Visible-infrared (RGB-IR) object detection leverages multimodal information to ensure reliable perception in complex environments. However, dynamic scenes pose significant challenges due to the frequent inconsistency between scene-level modality contributions and local spatial reliability. Furthermore, standard feature extraction progressively attenuates boundary-sensitive structural cues, and unified fusion strategies often fail to capture spatially varying cross-modal complementarity. To overcome these limitations, we propose a Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives. Specifically, a Geometric Boundary Enhancement Module (GBEM) embeds Sobel-based high-frequency priors into shallow dual-modal features via residual spatial modulation, preventing the loss of crucial localization cues during downsampling. In the deep semantic space, a Hybrid Dual-Perspective Adaptive Fusion Module (HDAM) employs an illumination-aware branch for global modality weighting and a spatial confidence-driven branch for local cross-modal rectification. A spatial gating mechanism then dynamically reconciles these macro-environmental and micro-signal features. Extensive experiments on M3FD, LLVIP, and DroneVehicle demonstrate the effectiveness of BDPNet. Compared with state-of-the-art methods, BDPNet improves mAP50-95 by 0.8% and 1.0% on M3FD and LLVIP, respectively, and improves mAP50 by 0.6% on DroneVehicle, while using substantially fewer parameters and lower computational cost.

Remote SensingVol. 18(18)
Fuzhou University (CN)
National Key Research and Development Program of China
Openalex Percentile: Top 13%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.