CSGE: A Modular Cross‐Spatial Guided Space Representation Enhancement for Multi‐View 3D Detection in Robotic System
ABSTRACT Multi‐view 3D detection plays a vital role in advancing the perception ability of robotic systems, with implicit methods achieving remarkable advancements, where the initial sparse object queries update iteratively through global attention with 2D image features and 3D position embedding. However, 2D features lack global spatial representation, while 3D position embedding exhibits weak spatial perception ability. Furthermore, misalignment between them in overlapping multi‐view space hampers the accurate capture of objects, thereby compromising detection accuracy. Our proposed CSGE is a plug‐and‐play module for general implicit methods, this module significantly improves the spatial representation of 2D features and the spatial perception ability of 3D position embedding. It addresses spatial misalignment in overlapping space of multi‐view images. By aligning 2D and 3D spatial features, CSGE strengthens the spatial understanding ability of robotic systems and broadens their range of applications. Comprehensive evaluations on the challenging nuScenes dataset demonstrate that CSGE consistently enhances detection performance in implicit methods using single‐frame and multi‐frame data. We qualitatively and quantitatively show that CSGE effectively bridges the feature gap between 2D and 3D space, underscoring the importance of addressing spatial feature discrepancies in general implicit methods.
Authors
- Shaoshan Liu (ORCID: https://orcid.org/0000-0002-5132-8351)
- Yefei Hou (ORCID: https://orcid.org/0009-0004-2624-9979)
- Jie Tang
- Bo Yu
Institutions
- Shenzhen Academy of Robotics (CN)
- South China University of Technology (CN)
Publication Details
- Journal
- Journal of Field Robotics
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1002/rob.70348
- Primary Topic
- Robotics and Sensor-Based Localization
- Type
- article
- Field-Weighted Citation Impact
- 0.00