Spatially Constrained 3D Scene Graphs for Scene-Grounded Subtask Generation

3D scene graphs connect spatial perception with high-level language reasoning, but unverified false-positive groundings can impair downstream subtask generation. We propose a feed-forward, spatially constrained framework that grounds task-relevant objects and generates position-aware subtasks without an iterative task–scene refinement loop. The framework performs Agglomerative Information Bottleneck clustering within room-specific spatial scopes and augments accumulated-feature similarity with a best-view raw image crop score. A Large Language Model (LLM) evaluates the combined confidence with spatial and geometric evidence to verify object candidates, and subsequent subtasks reference the retained object set. Controlled experiments across three reconstructed indoor scenes, Office, Apartment, and Cubicle, show that confidence-guided verification improves the unweighted scene-average Relaxed Accuracy F1 from 0.364 to 0.574 and pooled best-instance selection accuracy from 50.9% to 86.0% over 57 tasks. A separate open-set grounding evaluation shows higher precision with LLM-based verification, while the Office case study reveals a trade-off between fewer false positives and reduced retrieval coverage. The framework provides an intermediate scene-grounded subtask specification; residual grounding errors and omitted objects remain limitations for downstream use.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-28
DOI
https://doi.org/10.3390/electronics15194468
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Spatially Constrained 3D Scene Graphs for Scene-Grounded Subtask Generation

Soohwan Song, 류준상, Inseong Choi, Siwoo Lee
Electronics
Multimodal Machine Learning Applications
article

Spatially Constrained 3D Scene Graphs for Scene-Grounded Subtask Generation

Soohwan Song, 류준상, Inseong Choi, Siwoo Lee
article en

Abstract

3D scene graphs connect spatial perception with high-level language reasoning, but unverified false-positive groundings can impair downstream subtask generation. We propose a feed-forward, spatially constrained framework that grounds task-relevant objects and generates position-aware subtasks without an iterative task–scene refinement loop. The framework performs Agglomerative Information Bottleneck clustering within room-specific spatial scopes and augments accumulated-feature similarity with a best-view raw image crop score. A Large Language Model (LLM) evaluates the combined confidence with spatial and geometric evidence to verify object candidates, and subsequent subtasks reference the retained object set. Controlled experiments across three reconstructed indoor scenes, Office, Apartment, and Cubicle, show that confidence-guided verification improves the unweighted scene-average Relaxed Accuracy F1 from 0.364 to 0.574 and pooled best-instance selection accuracy from 50.9% to 86.0% over 57 tasks. A separate open-set grounding evaluation shows higher precision with LLM-based verification, while the Office case study reveals a trade-off between fewer false positives and reduced retrieval coverage. The framework provides an intermediate scene-grounded subtask specification; residual grounding errors and omitted objects remain limitations for downstream use.

ElectronicsVol. 15(19)
Dongguk University (KR)
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.