HADD: A hierarchy-aware disentangled diffusion framework with spatio-temporal denoising for 3D human pose estimation
Monocular 3D Human Pose Estimation (3D HPE) is a core task in computer vision, which provides critical technical support for applications including VR/AR, immersive human-computer interaction, digital human animation, as well as sports motion capture, athletic technique diagnosis, physical fitness assessment and sports rehabilitation training. However, existing methods still face three key limitations: diffusion-based methods lack explicit human anatomical prior constraints in the diffusion process; disentangled-based methods suffer from severe hierarchical error accumulation along the skeletal kinematic tree; and most mainstream methods fail to fully mine the fine-grained hierarchical spatio-temporal correlations of the human skeleton, leading to increased estimation error for high-hierarchy joints. To address these challenges, we propose a novel Hierarchy-Aware Disentangled Diffusion framework with Spatio-Temporal Denoising (HADD) for monocular 3D HPE, which deeply integrates disentanglement strategy with the diffusion model and embeds skeletal hierarchical information into the full diffusion pipeline. In the forward diffusion process, HADD disentangles 3D pose into bone length and bone direction features, and injects Gaussian noise into the two features separately. In the reverse denoising process, we design a Hierarchical Spatial and Temporal Denoising (HSTD) module to accurately model the hierarchical spatio-temporal relationships of the human skeleton and alleviate error amplification of high-hierarchy joints. Meanwhile, a hybrid loss function combining 3D disentanglement loss and 3D pose loss is constructed to realize dual supervision at the bone and joint levels. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets show that HADD achieves consistent performance improvement over state-of-the-art methods. On Human3.6M, it reduces average MPJPE by up to 12.9% compared with traditional disentangled-based methods, and outperforms mainstream diffusion-based probabilistic methods under lightweight inference configurations. On MPI-INF-3DHP, HADD achieves an MPJPE of 29.2 mm, a PCK of 98.5%, and an AUC of 78.1 under single-hypothesis inference, reaching new state-of-the-art performance on this dataset under the single-hypothesis inference setting with only ground-truth 2D joint sequences as input. This work provides an effective solution to the core limitations of existing 3D HPE methods, and also offers reliable technical support for fine-grained motion analysis and quantitative assessment in sports training and rehabilitation scenarios.
Authors
- Liming Zhang (ORCID: https://orcid.org/0000-0002-8451-4206)
- Yuxing Shi
- Minsi Liang
- Jiang Jinze
- Zhongheng Jian (ORCID: https://orcid.org/0009-0006-1300-6294)
- Hualing Zheng
Institutions
- Nanyang Technological University (SG)
- Hebei Petroleum University of Technology
- Sichuan Polytechnic University (CN)
- Inner Mongolia University of Technology (CN)
- Fujian Agriculture and Forestry University (CN)
- Fujian University of Technology (CN)
Publication Details
- Journal
- PLoS ONE
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1371/journal.pone.0359321
- Primary Topic
- Human Pose and Action Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00