Multi-Step Obstruction Reasoning for Target-Oriented Grasp Sequence Generation in Cluttered Scenes

Retrieving target objects in severe clutter requires multi-step reasoning to establish valid obstacle removal sequences. While recent Vision–Language Models (VLMs) have advanced instruction-driven clutter grasping, existing paradigms lack explicit construction of graph-constrained grasp sequences encompassing canonical trajectories and valid topological permutations to guide model fine-tuning and evaluation. In addition, comprehensive evaluation requires accounting for the full space of topologically valid clearing sequences while systematically disentangling high-level topological planning errors from low-level physical execution failures. To fulfill these requirements, we introduce a novel DAG-based obstruction reasoning framework coupled with an integrated diagnostic evaluation protocol. Specifically, we model scene-level physical dependencies as Directed Acyclic Graphs (DAGs), explicitly converting graph constraints into topologically feasible sequence permutations to drive VLM fine-tuning. For diagnostic evaluation, we establish a three-part offline protocol comprising Strict Exact Match (EM), Graph-Feasible Accuracy (GFA), and Multi-Reference Normalized Sequence Edit Distance (MR-NSED), paired with online simulation testing. Extensive experiments demonstrate that our topological fine-tuning significantly improves multi-path reasoning performance, outperforming strong foundation model baselines including GPT-4o, GPT-4o-mini, and Llama-3.2-90B-Vision-Instruct, while our diagnostic protocol provides a faithful mechanism to systematically isolate reasoning logic from manipulation mechanics in complex physical clutter.

Authors

Institutions

Publication Details

Journal
Robotics
Published
2026-09-22
DOI
https://doi.org/10.3390/robotics15100181
Primary Topic
Robot Manipulation and Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multi-Step Obstruction Reasoning for Target-Oriented Grasp Sequence Generation in Cluttered Scenes

Yuhua Zheng, Wei Song, Huixuan Yang, Shiqiang Zhu
Robotics
Robot Manipulation and Learning
article

Multi-Step Obstruction Reasoning for Target-Oriented Grasp Sequence Generation in Cluttered Scenes

Yuhua Zheng, Wei Song, Huixuan Yang, Shiqiang Zhu
article en

Abstract

Retrieving target objects in severe clutter requires multi-step reasoning to establish valid obstacle removal sequences. While recent Vision–Language Models (VLMs) have advanced instruction-driven clutter grasping, existing paradigms lack explicit construction of graph-constrained grasp sequences encompassing canonical trajectories and valid topological permutations to guide model fine-tuning and evaluation. In addition, comprehensive evaluation requires accounting for the full space of topologically valid clearing sequences while systematically disentangling high-level topological planning errors from low-level physical execution failures. To fulfill these requirements, we introduce a novel DAG-based obstruction reasoning framework coupled with an integrated diagnostic evaluation protocol. Specifically, we model scene-level physical dependencies as Directed Acyclic Graphs (DAGs), explicitly converting graph constraints into topologically feasible sequence permutations to drive VLM fine-tuning. For diagnostic evaluation, we establish a three-part offline protocol comprising Strict Exact Match (EM), Graph-Feasible Accuracy (GFA), and Multi-Reference Normalized Sequence Edit Distance (MR-NSED), paired with online simulation testing. Extensive experiments demonstrate that our topological fine-tuning significantly improves multi-path reasoning performance, outperforming strong foundation model baselines including GPT-4o, GPT-4o-mini, and Llama-3.2-90B-Vision-Instruct, while our diagnostic protocol provides a faithful mechanism to systematically isolate reasoning logic from manipulation mechanics in complex physical clutter.

RoboticsVol. 15(10)
Chinese Academy of Sciences (CN), Institute of Computing Technology (CN), Robotics Research (United States) (US), University of Chinese Academy of Sciences (CN), Zhejiang University (CN)
Openalex Percentile: Top 15%
Robot Manipulation and Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.