Multi-View Semantic–Geometric Perception for Post-Disaster Flexible Suction with Metric 3D Scene Reconstruction
Post-disaster flexible suction requires perception of obstacle semantics and metric 3D geometry in cluttered, partially occluded workspaces. However, image-plane masks alone cannot provide physical height, surface slope, or obstacle clearance, while geometry without semantics cannot distinguish obstacle regions from obstacle-free sludge regions. This study proposes a multi-view semantic-geometric perception framework using synchronized arm-mounted RGB images. YOLO-LiteSeg provides obstacle instance masks, while Visual Geometry Grounded Transformer (VGGT) provides depth and camera parameters for multi-view reconstruction; checkerboard-based metric alignment converts reconstruction to physical scale. Projection-based 2D-to-3D semantic label transfer produces an obstacle-highlighted metric 3D scene representation for constraint-based suction-point selection. Across nine simulated conditions, YOLO-LiteSeg reduced parameters from 27.24 M to 2.61 M and giga floating-point operations from 52.37 to 5.47, while achieving 99.02% mean average precision at 0.50 intersection over union and 84.36% across thresholds from 0.50 to 0.95, with 4.56 ms inference. Reconstructed steepest-slope height profiles showed mean absolute errors below 2 cm. Enhanced and redundant camera configurations yielded a mean projected camera-position error of 4.02 cm and a root mean square error of 4.53 cm. The resulting 3D representation generated suction-point references satisfying predefined suction-depth and obstacle-clearance criteria, providing metrically interpretable spatial information for post-disaster flexible suction.
Authors
- Hong Weng
- Yuanlong Yu
Institutions
- Fuzhou University (CN)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-10-09
- DOI
- https://doi.org/10.3390/app162010003
- Primary Topic
- Advanced Vision and Imaging
- Type
- article
- Field-Weighted Citation Impact
- 0.00