Skeleton-guided geometry-aware auto-labeling for robotic grasping: From 4-DoF learning to 6-DoF execution
Robotic grasping relies on large manually annotated datasets, which are costly to create. This work introduces a novel algorithm for generating grasp keypoints based on straight skeletons, which identify high-potential regions for stable and feasible grasp pose estimation. A fully automated pipeline has been developed to perform grasp auto-labeling without human intervention, facilitating large-scale annotation with consistent quality. Furthermore, a new architecture for 4-DoF grasping, named the Skeleton-based Generative Grasp Convolutional Neural Network (SkelGG-CNN), is introduced, which incorporates skeleton-based guidance during training to directly learn geometric grasping features. Finally, the developed auto-labeling is extended to 6-DoF grasping, applied directly to point clouds via a generalized quasi-3D skeletonization algorithm. Evaluations demonstrate that SkelGG-CNN achieves 95.42%, 96.55%, and 97.1% accuracy on Jacquard V1, Jacquard V2, and Cornell benchmarks, respectively, using 4-DoF auto-generated labels. On GraspNet, the 6-DoF auto-labeling algorithm yields an average precision of 18.33 on the novel set, comparable to state-of-the-art model-based approaches. In simulation, the generalized 6-DoF auto-labeling method yields success rates of 98.4% on the Dex-Net test set and 89.1% on EGAD!. Real-world experiments on a Delta Parallel robot confirm 97.4% success on single household objects and 95.0% in cluttered scenes, outperforming state-of-the-art methods without domain adaptation. Experiments show that the framework’s 4-DoF pipeline is adaptable across environments, generating grasp annotations that approximate human-level precision, enabling efficient training of lightweight models like SkelGG-CNN on small datasets for real-world deployment. Additionally, the generalized 6-DoF auto-labeling algorithm extends this automation to point clouds, providing scalable, geometry-aware annotations for cluttered 3D scenes. Source code and videos are available at https://github.com/Farbod82/SkelGGCNN .
Authors
- Ahmad Kalhor (ORCID: https://orcid.org/0000-0001-6657-6705)
- Mehdi Tale Masouleh (ORCID: https://orcid.org/0000-0001-8249-1998)
- Ali Sabzejou
- Farbod Azimmohseni
Institutions
- University of Tehran (IR)
Publication Details
- Journal
- The International Journal of Robotics Research
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1177/02783649261477772
- Primary Topic
- Robot Manipulation and Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00