BMF-DETR: Pseudo-Depth-Guided Bidirectional Multi-Strategy Fusion for End-to-End Object Detection
Transformer-based detectors model long-range context effectively, yet their representations remain dominated by RGB appearance and can become unreliable in cluttered, occluded, or crowded scenes. We present BMF-DETR, a pseudo-depth-guided detector that introduces RGB-derived geometric structure without requiring a depth sensor. DA3Mono-Large from Depth Anything 3 generates spatially aligned pseudo-depth maps offline, while two ResNet-50 streams encode appearance and relative geometry. Bidirectional cross-modal attention (BCMA) establishes two-way correspondence, and multi-strategy fusion (MSF) combines the streams through global calibration, channel allocation, and spatial gating before squeeze-and-excitation (SE) recalibration. On the fixed validation/evaluation split of the 2024 Roboflow-curated PASCAL VOC derivative, the complete model reaches 60.80 AP, compared with 52.80 AP for a capacity-matched dual-RGB control. BMF-DETR obtains 49.30 AP on COCO 2017. Its detector contains 58 M parameters and requires 103 GFLOPs; these figures exclude offline pseudo-depth generation. A shared-low-level variant retains 60.10 AP with 50 M parameters and 87 GFLOPs. The results show that pseudo-depth can serve as a useful auxiliary representation when its contribution is separated from capacity effects and evaluated under controlled fusion settings.
Authors
- Chunlai Yang (ORCID: https://orcid.org/0000-0003-4604-6592)
- Junhao Wen (ORCID: https://orcid.org/0000-0002-6561-560X)
- Hai Wang (ORCID: https://orcid.org/0000-0001-7752-1013)
- Jiale Gu
- Kamara Kekele Adnan Fayçal
Institutions
- Anhui Polytechnic University (CN)
Publication Details
- Journal
- AI
- Published
- 2026-09-21
- DOI
- https://doi.org/10.3390/ai7090384
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00