Hierarchical Manipulation Skill Learning for Long-Horizon Robotic Manipulation Based on Residual Finite Scalar Quantization

Leveraging discrete sequence prediction for manipulation skill learning has demonstrated immense potential for complex long-horizon tasks. However, guaranteeing precise and robust execution under complex scenarios remains a core challenge for efficient autonomous robotic operation. This demands execution policies capable of generating precise and efficient action sequences based on perceptual feedback. To address the loss of fine-grained action details induced by action discretization and the difficulty in preserving long-horizon temporal dependencies in long-horizon tasks, this work proposed a hierarchical manipulation skill learning method based on residual finite scalar quantization (RFSQ). At the action representation level, a robust RFSQ with an adaptive linear scaling mechanism was introduced. It effectively mitigated residual decay in multi-level action quantization and realized accurate discrete reconstruction of continuous action sequences. At the action sequence generation level, a hybrid hierarchical architecture integrating Mamba2 and Transformer was constructed. Mamba2 efficiently captured temporal dependencies among actions, while Transformer precisely modeled fine-grained residual dependencies across action quantization hierarchies. A multimodal perception fusion module was further incorporated to strengthen instruction-semantic alignment on language-instructed tasks. A series of experiments were conducted on the LIBERO and Meta-World simulation environments. Experimental results demonstrated that the proposed method achieved outstanding performance in multi-task and long-horizon task learning.

Authors

Institutions

Publication Details

Journal
Sensors
Published
2026-10-06
DOI
https://doi.org/10.3390/s26196310
Primary Topic
Robot Manipulation and Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Hierarchical Manipulation Skill Learning for Long-Horizon Robotic Manipulation Based on Residual Finite Scalar Quantization

Yongfeng Rong, Jiahui Guo, Guanghui Ma, Huaidong Zhou et al.
Sensors
Robot Manipulation and Learning
article

Hierarchical Manipulation Skill Learning for Long-Horizon Robotic Manipulation Based on Residual Finite Scalar Quantization

Yongfeng Rong, Jiahui Guo, Guanghui Ma, Huaidong Zhou, Xinhua Tang
article en

Abstract

Leveraging discrete sequence prediction for manipulation skill learning has demonstrated immense potential for complex long-horizon tasks. However, guaranteeing precise and robust execution under complex scenarios remains a core challenge for efficient autonomous robotic operation. This demands execution policies capable of generating precise and efficient action sequences based on perceptual feedback. To address the loss of fine-grained action details induced by action discretization and the difficulty in preserving long-horizon temporal dependencies in long-horizon tasks, this work proposed a hierarchical manipulation skill learning method based on residual finite scalar quantization (RFSQ). At the action representation level, a robust RFSQ with an adaptive linear scaling mechanism was introduced. It effectively mitigated residual decay in multi-level action quantization and realized accurate discrete reconstruction of continuous action sequences. At the action sequence generation level, a hybrid hierarchical architecture integrating Mamba2 and Transformer was constructed. Mamba2 efficiently captured temporal dependencies among actions, while Transformer precisely modeled fine-grained residual dependencies across action quantization hierarchies. A multimodal perception fusion module was further incorporated to strengthen instruction-semantic alignment on language-instructed tasks. A series of experiments were conducted on the LIBERO and Meta-World simulation environments. Experimental results demonstrated that the proposed method achieved outstanding performance in multi-task and long-horizon task learning.

SensorsVol. 26(19)
Northwestern Polytechnical University (CN), Shenzhen University (CN), Anhui Polytechnic University (CN), Tsinghua University (CN)
Openalex Percentile: Top 16%
Robot Manipulation and Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Hierarchical Manipulation Skill Learning for Long-Horizon Robotic Manipulation Based on Residual Finite Scalar Quantization — Yongfeng Rong, Jiahui Guo, et al. · Sensors (2026) | TGRS Research Map | TGRS