Spatially Constrained 3D Scene Graphs for Scene-Grounded Subtask Generation
3D scene graphs connect spatial perception with high-level language reasoning, but unverified false-positive groundings can impair downstream subtask generation. We propose a feed-forward, spatially constrained framework that grounds task-relevant objects and generates position-aware subtasks without an iterative task–scene refinement loop. The framework performs Agglomerative Information Bottleneck clustering within room-specific spatial scopes and augments accumulated-feature similarity with a best-view raw image crop score. A Large Language Model (LLM) evaluates the combined confidence with spatial and geometric evidence to verify object candidates, and subsequent subtasks reference the retained object set. Controlled experiments across three reconstructed indoor scenes, Office, Apartment, and Cubicle, show that confidence-guided verification improves the unweighted scene-average Relaxed Accuracy F1 from 0.364 to 0.574 and pooled best-instance selection accuracy from 50.9% to 86.0% over 57 tasks. A separate open-set grounding evaluation shows higher precision with LLM-based verification, while the Office case study reveals a trade-off between fewer false positives and reduced retrieval coverage. The framework provides an intermediate scene-grounded subtask specification; residual grounding errors and omitted objects remain limitations for downstream use.
Authors
- Soohwan Song (ORCID: https://orcid.org/0000-0003-2145-1161)
- 류준상
- Inseong Choi
- Siwoo Lee
Institutions
- Dongguk University (KR)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-28
- DOI
- https://doi.org/10.3390/electronics15194468
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00