SE-Former: Skeleton-Enhanced Learning Framework for 3D Human Motion Prediction

The key challenge in 3D human motion prediction lies in the inaccurate modeling of spatial and temporal dependencies in motion sequences. The latent feature relationships within the data remain underexplored, particularly with incomplete data (missing joints). Motivated by the coordinated motion of multiple skeletons in the human body, this paper presents SE-Former, a skeleton-enhanced Transformer model designed to advance the feature representation of human motion. It leverages the spatial-temporal information of motion sequences, and mines the latent skeletal spatial information to adapt to the intrinsic properties of the data. Specifically, SE-Former introduces virtual joints with noise and simulates realistic joint combinations, which are then integrated with native skeletal features via a graph convolutional network-attention encoder. Considering the long-term prediction is easily hindered by the model tendency to converge to average poses, we draw inspiration from trend modelling techniques in the biomechanics domain. A trend learning module is integrated to capture long-term motion trends and residual dynamics. It effectively bridges the encoder and decoder, ensuring that predictions maintain dynamic variability, so as to avoid the convergence toward average poses and improve the fidelity of motion predictions. Extensive experiments are conducted on Human3.6M, CMU-MoCap and 3DPW datasets, where our SE-Former shows significant advantages over the other state-of-the-art work. Code is available at https://github.com/Logan-007L/SEFormer .

Authors

Institutions

Publication Details

Journal
ACM Transactions on Intelligent Systems and Technology
Published
2026-10-06
DOI
https://doi.org/10.1145/3856803
Primary Topic
Human Motion and Animation
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

SE-Former: Skeleton-Enhanced Learning Framework for 3D Human Motion Prediction

Yushan Pan, Shu Liu, Luo Zheng, Jiaheng Wang
ACM Transactions on Intelligent Systems and Technology
Human Motion and Animation
article

SE-Former: Skeleton-Enhanced Learning Framework for 3D Human Motion Prediction

Yushan Pan, Shu Liu, Luo Zheng, Jiaheng Wang
article en

Abstract

The key challenge in 3D human motion prediction lies in the inaccurate modeling of spatial and temporal dependencies in motion sequences. The latent feature relationships within the data remain underexplored, particularly with incomplete data (missing joints). Motivated by the coordinated motion of multiple skeletons in the human body, this paper presents SE-Former, a skeleton-enhanced Transformer model designed to advance the feature representation of human motion. It leverages the spatial-temporal information of motion sequences, and mines the latent skeletal spatial information to adapt to the intrinsic properties of the data. Specifically, SE-Former introduces virtual joints with noise and simulates realistic joint combinations, which are then integrated with native skeletal features via a graph convolutional network-attention encoder. Considering the long-term prediction is easily hindered by the model tendency to converge to average poses, we draw inspiration from trend modelling techniques in the biomechanics domain. A trend learning module is integrated to capture long-term motion trends and residual dynamics. It effectively bridges the encoder and decoder, ensuring that predictions maintain dynamic variability, so as to avoid the convergence toward average poses and improve the fidelity of motion predictions. Extensive experiments are conducted on Human3.6M, CMU-MoCap and 3DPW datasets, where our SE-Former shows significant advantages over the other state-of-the-art work. Code is available at https://github.com/Logan-007L/SEFormer .

ACM Transactions on Intelligent Systems and Technology
Central South University (CN), Alibaba Group (China) (CN), Xi’an Jiaotong-Liverpool University (CN)
Openalex Percentile: Top 16%
Human Motion and Animation
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

SE-Former: Skeleton-Enhanced Learning Framework for 3D Human Motion Prediction — Yushan Pan, Shu Liu, et al. · ACM Transactions on Intelligent Systems and Technology (2026) | TGRS Research Map | TGRS