Spatio-temporal collaborative enhancement for 3D human pose estimation in complex dynamic scenes

To address the challenges of pose “jumping” caused by insufficient spatio-temporal feature coupling in existing diffusion-based 3D human pose estimation methods under dynamic scenes, as well as their suboptimal accuracy and robustness in complex scenarios, this paper proposes an enhanced high-reliability human pose estimation method based on the FinePose model. Firstly, a spatio-temporal noise module is designed following the generation paradigm of spatio-temporal noise in stochastic processes. This module integrates one-dimensional average pooling with the skeleton adjacency matrix to generate and inject spatio-temporal noise into the diffusion process, thereby eliminating pose jumps in dynamic scenes. Secondly, considering the non-Euclidean geometry of human skeletons, a single-layer Graph Convolutional Network (GCN) with residual connections is incorporated to construct a human skeletal topology graph. This enhances the spatial relationships between joints and improves estimation performance in self-occlusion scenarios. Finally, a multi-scale position encoding is designed and cross-domain fusion with GCN spatial features is performed. This combination overcomes the Transformer’s limitations in scale perception for time-series prediction and enhances the model’s adaptability to challenging imaging conditions. Experiments on the Human3.6M and MPI-INF-3DHP datasets demonstrate that compared to 13 mainstream algorithms, the proposed method achieves an MPJPE as low as 29.6 mm, which is 2.3 mm lower than the original FinePose model. Significant improvements in MPJPE accuracy are also observed in complex scenarios such as dynamic self-occlusion and static fine-grained movements, effectively achieving a synergistic enhancement of both accuracy and robustness. https://github.com/xiao-qiang-coal-mine/SGMPOSE.git .

Authors

Institutions

Publication Details

Journal
Discover Applied Sciences
Published
2026-09-25
DOI
https://doi.org/10.1007/s42452-026-09596-9
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Spatio-temporal collaborative enhancement for 3D human pose estimation in complex dynamic scenes

Weiqiang Fan, Xiaoyu Li, Meng Lv
Discover Applied Sciences
Human Pose and Action Recognition
article

Spatio-temporal collaborative enhancement for 3D human pose estimation in complex dynamic scenes

Weiqiang Fan, Xiaoyu Li, Meng Lv
article en

Abstract

To address the challenges of pose “jumping” caused by insufficient spatio-temporal feature coupling in existing diffusion-based 3D human pose estimation methods under dynamic scenes, as well as their suboptimal accuracy and robustness in complex scenarios, this paper proposes an enhanced high-reliability human pose estimation method based on the FinePose model. Firstly, a spatio-temporal noise module is designed following the generation paradigm of spatio-temporal noise in stochastic processes. This module integrates one-dimensional average pooling with the skeleton adjacency matrix to generate and inject spatio-temporal noise into the diffusion process, thereby eliminating pose jumps in dynamic scenes. Secondly, considering the non-Euclidean geometry of human skeletons, a single-layer Graph Convolutional Network (GCN) with residual connections is incorporated to construct a human skeletal topology graph. This enhances the spatial relationships between joints and improves estimation performance in self-occlusion scenarios. Finally, a multi-scale position encoding is designed and cross-domain fusion with GCN spatial features is performed. This combination overcomes the Transformer’s limitations in scale perception for time-series prediction and enhances the model’s adaptability to challenging imaging conditions. Experiments on the Human3.6M and MPI-INF-3DHP datasets demonstrate that compared to 13 mainstream algorithms, the proposed method achieves an MPJPE as low as 29.6 mm, which is 2.3 mm lower than the original FinePose model. Significant improvements in MPJPE accuracy are also observed in complex scenarios such as dynamic self-occlusion and static fine-grained movements, effectively achieving a synergistic enhancement of both accuracy and robustness. https://github.com/xiao-qiang-coal-mine/SGMPOSE.git .

Discover Applied Sciences
Inner Mongolia University (CN)
Openalex Percentile: Top 14%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.