An Adaptive IMU–Visual Multimodal Fusion System for Real-Time Exercise Recognition and Movement Quality Assessment

Exercise recognition and movement quality assessment remain challenging in supervised exercise training, particularly under viewpoint changes and self-occlusion. Vision-based methods provide spatial posture information but are susceptible to keypoint loss, whereas inertial sensing is less affected by occlusion but provides limited information about global posture geometry. This study presents a dual-stream prototype that combines a nine-axis inertial measurement unit (IMU) with vision-based pose estimation. A 1DCNN-LSTM branch models inertial dynamics, a custom keypoint temporal branch models normalized pose sequences, and a confidence-gated rule adjusts their contributions according to visual keypoint reliability. Owing to the absence of a public synchronized multi-view IMU–vision exercise dataset with the required protocol, we constructed IMV-Exercise, comprising 10 participants, three exercises, and 900 repetition-level samples with side-, front-, and posterior-view recordings. The system achieved 96.0% exercise recognition accuracy under leave-one-subject-out cross-validation. In a separate viewpoint-specific evaluation, the fused output achieved 91.2% action-window accuracy under posterior viewing. Across 50 online trials, the reported recognition accuracy was 96.0%, the mean end-to-end latency was 195 ms, and the recorded maximum was below 210 ms. These results establish feasibility within the studied cohort and exercises; broader generalization and feedback effectiveness require larger, independently controlled evaluations.

Authors

Institutions

Publication Details

Journal
Sensors
Published
2026-09-15
DOI
https://doi.org/10.3390/s26185838
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

An Adaptive IMU–Visual Multimodal Fusion System for Real-Time Exercise Recognition and Movement Quality Assessment

Zhaoyang Gu, Ruopeng Yang, Chaoyang Li, Bo Huang et al.
Sensors
Human Pose and Action Recognition
article

An Adaptive IMU–Visual Multimodal Fusion System for Real-Time Exercise Recognition and Movement Quality Assessment

Zhaoyang Gu, Ruopeng Yang, Chaoyang Li, Bo Huang, Kaige Jiao, Yongqi Wen, Yihao Zhong, Chen He, Yu Tao, Yongqi Shi, Dongxu Dai
article en

Abstract

Exercise recognition and movement quality assessment remain challenging in supervised exercise training, particularly under viewpoint changes and self-occlusion. Vision-based methods provide spatial posture information but are susceptible to keypoint loss, whereas inertial sensing is less affected by occlusion but provides limited information about global posture geometry. This study presents a dual-stream prototype that combines a nine-axis inertial measurement unit (IMU) with vision-based pose estimation. A 1DCNN-LSTM branch models inertial dynamics, a custom keypoint temporal branch models normalized pose sequences, and a confidence-gated rule adjusts their contributions according to visual keypoint reliability. Owing to the absence of a public synchronized multi-view IMU–vision exercise dataset with the required protocol, we constructed IMV-Exercise, comprising 10 participants, three exercises, and 900 repetition-level samples with side-, front-, and posterior-view recordings. The system achieved 96.0% exercise recognition accuracy under leave-one-subject-out cross-validation. In a separate viewpoint-specific evaluation, the fused output achieved 91.2% action-window accuracy under posterior viewing. Across 50 online trials, the reported recognition accuracy was 96.0%, the mean end-to-end latency was 195 ms, and the recorded maximum was below 210 ms. These results establish feasibility within the studied cohort and exercises; broader generalization and feedback effectiveness require larger, independently controlled evaluations.

SensorsVol. 26(18)
PLA Information Engineering University (CN), National University of Defense Technology (CN)
Openalex Percentile: Top 13%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.