Spatio-temporal collaborative enhancement for 3D human pose estimation in complex dynamic scenes
To address the challenges of pose “jumping” caused by insufficient spatio-temporal feature coupling in existing diffusion-based 3D human pose estimation methods under dynamic scenes, as well as their suboptimal accuracy and robustness in complex scenarios, this paper proposes an enhanced high-reliability human pose estimation method based on the FinePose model. Firstly, a spatio-temporal noise module is designed following the generation paradigm of spatio-temporal noise in stochastic processes. This module integrates one-dimensional average pooling with the skeleton adjacency matrix to generate and inject spatio-temporal noise into the diffusion process, thereby eliminating pose jumps in dynamic scenes. Secondly, considering the non-Euclidean geometry of human skeletons, a single-layer Graph Convolutional Network (GCN) with residual connections is incorporated to construct a human skeletal topology graph. This enhances the spatial relationships between joints and improves estimation performance in self-occlusion scenarios. Finally, a multi-scale position encoding is designed and cross-domain fusion with GCN spatial features is performed. This combination overcomes the Transformer’s limitations in scale perception for time-series prediction and enhances the model’s adaptability to challenging imaging conditions. Experiments on the Human3.6M and MPI-INF-3DHP datasets demonstrate that compared to 13 mainstream algorithms, the proposed method achieves an MPJPE as low as 29.6 mm, which is 2.3 mm lower than the original FinePose model. Significant improvements in MPJPE accuracy are also observed in complex scenarios such as dynamic self-occlusion and static fine-grained movements, effectively achieving a synergistic enhancement of both accuracy and robustness. https://github.com/xiao-qiang-coal-mine/SGMPOSE.git .
Authors
- Weiqiang Fan (ORCID: https://orcid.org/0000-0002-4799-5223)
- Xiaoyu Li (ORCID: https://orcid.org/0000-0002-0995-5095)
- Meng Lv (ORCID: https://orcid.org/0000-0002-2804-4342)
Institutions
- Inner Mongolia University (CN)
Publication Details
- Journal
- Discover Applied Sciences
- Published
- 2026-09-25
- DOI
- https://doi.org/10.1007/s42452-026-09596-9
- Primary Topic
- Human Pose and Action Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00