HADD: A hierarchy-aware disentangled diffusion framework with spatio-temporal denoising for 3D human pose estimation

Monocular 3D Human Pose Estimation (3D HPE) is a core task in computer vision, which provides critical technical support for applications including VR/AR, immersive human-computer interaction, digital human animation, as well as sports motion capture, athletic technique diagnosis, physical fitness assessment and sports rehabilitation training. However, existing methods still face three key limitations: diffusion-based methods lack explicit human anatomical prior constraints in the diffusion process; disentangled-based methods suffer from severe hierarchical error accumulation along the skeletal kinematic tree; and most mainstream methods fail to fully mine the fine-grained hierarchical spatio-temporal correlations of the human skeleton, leading to increased estimation error for high-hierarchy joints. To address these challenges, we propose a novel Hierarchy-Aware Disentangled Diffusion framework with Spatio-Temporal Denoising (HADD) for monocular 3D HPE, which deeply integrates disentanglement strategy with the diffusion model and embeds skeletal hierarchical information into the full diffusion pipeline. In the forward diffusion process, HADD disentangles 3D pose into bone length and bone direction features, and injects Gaussian noise into the two features separately. In the reverse denoising process, we design a Hierarchical Spatial and Temporal Denoising (HSTD) module to accurately model the hierarchical spatio-temporal relationships of the human skeleton and alleviate error amplification of high-hierarchy joints. Meanwhile, a hybrid loss function combining 3D disentanglement loss and 3D pose loss is constructed to realize dual supervision at the bone and joint levels. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets show that HADD achieves consistent performance improvement over state-of-the-art methods. On Human3.6M, it reduces average MPJPE by up to 12.9% compared with traditional disentangled-based methods, and outperforms mainstream diffusion-based probabilistic methods under lightweight inference configurations. On MPI-INF-3DHP, HADD achieves an MPJPE of 29.2 mm, a PCK of 98.5%, and an AUC of 78.1 under single-hypothesis inference, reaching new state-of-the-art performance on this dataset under the single-hypothesis inference setting with only ground-truth 2D joint sequences as input. This work provides an effective solution to the core limitations of existing 3D HPE methods, and also offers reliable technical support for fine-grained motion analysis and quantitative assessment in sports training and rehabilitation scenarios.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-28
DOI
https://doi.org/10.1371/journal.pone.0359321
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

HADD: A hierarchy-aware disentangled diffusion framework with spatio-temporal denoising for 3D human pose estimation

Liming Zhang, Yuxing Shi, Minsi Liang, Jiang Jinze et al.
PLoS ONE
Human Pose and Action Recognition
article

HADD: A hierarchy-aware disentangled diffusion framework with spatio-temporal denoising for 3D human pose estimation

Liming Zhang, Yuxing Shi, Minsi Liang, Jiang Jinze, Zhongheng Jian, Hualing Zheng
article en

Abstract

Monocular 3D Human Pose Estimation (3D HPE) is a core task in computer vision, which provides critical technical support for applications including VR/AR, immersive human-computer interaction, digital human animation, as well as sports motion capture, athletic technique diagnosis, physical fitness assessment and sports rehabilitation training. However, existing methods still face three key limitations: diffusion-based methods lack explicit human anatomical prior constraints in the diffusion process; disentangled-based methods suffer from severe hierarchical error accumulation along the skeletal kinematic tree; and most mainstream methods fail to fully mine the fine-grained hierarchical spatio-temporal correlations of the human skeleton, leading to increased estimation error for high-hierarchy joints. To address these challenges, we propose a novel Hierarchy-Aware Disentangled Diffusion framework with Spatio-Temporal Denoising (HADD) for monocular 3D HPE, which deeply integrates disentanglement strategy with the diffusion model and embeds skeletal hierarchical information into the full diffusion pipeline. In the forward diffusion process, HADD disentangles 3D pose into bone length and bone direction features, and injects Gaussian noise into the two features separately. In the reverse denoising process, we design a Hierarchical Spatial and Temporal Denoising (HSTD) module to accurately model the hierarchical spatio-temporal relationships of the human skeleton and alleviate error amplification of high-hierarchy joints. Meanwhile, a hybrid loss function combining 3D disentanglement loss and 3D pose loss is constructed to realize dual supervision at the bone and joint levels. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets show that HADD achieves consistent performance improvement over state-of-the-art methods. On Human3.6M, it reduces average MPJPE by up to 12.9% compared with traditional disentangled-based methods, and outperforms mainstream diffusion-based probabilistic methods under lightweight inference configurations. On MPI-INF-3DHP, HADD achieves an MPJPE of 29.2 mm, a PCK of 98.5%, and an AUC of 78.1 under single-hypothesis inference, reaching new state-of-the-art performance on this dataset under the single-hypothesis inference setting with only ground-truth 2D joint sequences as input. This work provides an effective solution to the core limitations of existing 3D HPE methods, and also offers reliable technical support for fine-grained motion analysis and quantitative assessment in sports training and rehabilitation scenarios.

PLoS ONEVol. 21(9)
Nanyang Technological University (SG), Hebei Petroleum University of Technology, Sichuan Polytechnic University (CN), Inner Mongolia University of Technology (CN), Fujian Agriculture and Forestry University (CN), Fujian University of Technology (CN)
Openalex Percentile: Top 14%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.