Detection-Guided Keypoint Estimation for Humanoid Robots from Video Frames

Reliable estimation of humanoid robot body configuration from monocular RGB frames is useful for external monitoring in traffic, logistics, and service-robotics environments. However, direct transfer of human pose models is challenged by differences in body proportions, rigid surface appearance, joint morphology, and local texture, while robot-specific annotated data are often limited. This study presents a detection-guided framework for full-body 2D keypoint estimation from monocular video frames. A robot-specific detector first localises the target, and the detected box is expanded and normalised into a local region of interest (ROI) for keypoint recovery. Rather than treating the task as direct coordinate regression, the proposed integration uses ResNet18 feature extraction, convolutional block attention module (CBAM) refinement, heatmap-based keypoint representation, differentiable spatial to numerical transform (DSNT)-based continuous coordinate decoding, and skeleton-aware regularisation. A compact 13-keypoint annotation protocol is defined to describe the head, shoulders, elbows, hands, hips, knees, and feet. On the held-out test set, the proposed model reduces mean per-joint position error (MPJPE) from 27.98 px to 15.49 px relative to the ResNet18 direct-regression baseline and improves object keypoint similarity (OKS) from 0.80 to 0.96. Under the same 13-keypoint protocol, the proposed model also achieves a lower MPJPE than fine-tuned YOLOv8-Pose with detector-ROI input (18.85 px) and HRNet-W18 (30.82 px). Direct transfer of COCO-pretrained YOLOv8-Pose performs substantially worse, with an MPJPE of 201.50 px and an OKS of 0.28. These results support the effectiveness of robot-specific local spatial modelling and continuous coordinate recovery for monocular humanoid robot keypoint estimation under the evaluated limited-data setting.

Authors

Institutions

Publication Details

Journal
Sensors
Published
2026-09-11
DOI
https://doi.org/10.3390/s26185784
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Detection-Guided Keypoint Estimation for Humanoid Robots from Video Frames

Zhihuo Xu, Yuexia Wang, Xuan Lou
Sensors
Human Pose and Action Recognition
article

Detection-Guided Keypoint Estimation for Humanoid Robots from Video Frames

Zhihuo Xu, Yuexia Wang, Xuan Lou
article en

Abstract

Reliable estimation of humanoid robot body configuration from monocular RGB frames is useful for external monitoring in traffic, logistics, and service-robotics environments. However, direct transfer of human pose models is challenged by differences in body proportions, rigid surface appearance, joint morphology, and local texture, while robot-specific annotated data are often limited. This study presents a detection-guided framework for full-body 2D keypoint estimation from monocular video frames. A robot-specific detector first localises the target, and the detected box is expanded and normalised into a local region of interest (ROI) for keypoint recovery. Rather than treating the task as direct coordinate regression, the proposed integration uses ResNet18 feature extraction, convolutional block attention module (CBAM) refinement, heatmap-based keypoint representation, differentiable spatial to numerical transform (DSNT)-based continuous coordinate decoding, and skeleton-aware regularisation. A compact 13-keypoint annotation protocol is defined to describe the head, shoulders, elbows, hands, hips, knees, and feet. On the held-out test set, the proposed model reduces mean per-joint position error (MPJPE) from 27.98 px to 15.49 px relative to the ResNet18 direct-regression baseline and improves object keypoint similarity (OKS) from 0.80 to 0.96. Under the same 13-keypoint protocol, the proposed model also achieves a lower MPJPE than fine-tuned YOLOv8-Pose with detector-ROI input (18.85 px) and HRNet-W18 (30.82 px). Direct transfer of COCO-pretrained YOLOv8-Pose performs substantially worse, with an MPJPE of 201.50 px and an OKS of 0.28. These results support the effectiveness of robot-specific local spatial modelling and continuous coordinate recovery for monocular humanoid robot keypoint estimation under the evaluated limited-data setting.

SensorsVol. 26(18)
Nantong University (CN)
National Natural Science Foundation of China
Openalex Percentile: Top 13%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Detection-Guided Keypoint Estimation for Humanoid Robots from Video Frames — Zhihuo Xu, Yuexia Wang, et al. · Sensors (2026) | TGRS Research Map | TGRS