Multimodal deep learning architecture for gym exercise recognition and real-time posture error detection using LSTM and RAG-based expert chatbot

The growing popularity of home fitness has led to an increase in home-based fitness training. Still, the lack of supervision by a trained professional has resulted in improper form, thereby increasing the risk of injury during exercise. Automated computer vision systems have also been developed to monitor exercise, but these systems rely on rigid joint-angle thresholds or require large annotated datasets of improper exercise. This paper discusses an artificial intelligence (AI) fitness coaching system that uses real-time pose-based analysis of exercise and a chatbot with expert knowledge for interactive communication with the user. The system uses MediaPipe for real-time two-dimensional pose estimation from the webcam stream, then processes the resulting landmark sequences with temporal deep learning models. A Long Short-Term Memory (LSTM) network classifies the exercise being performed, and a set of exercise-specific LSTM-based encoder-decoder networks, trained only on correct repetitions, evaluate form correctness through a reconstruction-error anomaly-detection criterion. A Retrieval-Augmented Generation (RAG) component built on the Gemma-2B-IT language model provides natural-language explanations and contextual coaching. Evaluation on a custom dataset of 9,414 videos covering three dumbbell exercises (bicep curl, shoulder press, lateral side raise) yields 95.9% classification accuracy, with a Bi-LSTM + Attention variant reaching 96.8%. The encoder-decoder reconstruction error separates correct and incorrect forms by an order of magnitude (mean ratio 10×–15×). The Chatbot achieves a perplexity of 6.8 and a mean human-rated helpfulness of 4.4/5 across 15 evaluators. The proposed multimodal architecture demonstrates a practical pathway toward intelligent home-fitness coaching.

Authors

Institutions

Publication Details

Journal
Journal of Health Population and Nutrition
Published
2026-10-07
DOI
https://doi.org/10.1186/s41043-026-01458-9
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Multimodal deep learning architecture for gym exercise recognition and real-time posture error detection using LSTM and RAG-based expert chatbot

Sourav Samaddar, B. Indira, Ruben Johnson Robert Jeremiah, J. Omana
Journal of Health Population and Nutrition
Human Pose and Action Recognition
article

Multimodal deep learning architecture for gym exercise recognition and real-time posture error detection using LSTM and RAG-based expert chatbot

Sourav Samaddar, B. Indira, Ruben Johnson Robert Jeremiah, J. Omana
article en

Abstract

The growing popularity of home fitness has led to an increase in home-based fitness training. Still, the lack of supervision by a trained professional has resulted in improper form, thereby increasing the risk of injury during exercise. Automated computer vision systems have also been developed to monitor exercise, but these systems rely on rigid joint-angle thresholds or require large annotated datasets of improper exercise. This paper discusses an artificial intelligence (AI) fitness coaching system that uses real-time pose-based analysis of exercise and a chatbot with expert knowledge for interactive communication with the user. The system uses MediaPipe for real-time two-dimensional pose estimation from the webcam stream, then processes the resulting landmark sequences with temporal deep learning models. A Long Short-Term Memory (LSTM) network classifies the exercise being performed, and a set of exercise-specific LSTM-based encoder-decoder networks, trained only on correct repetitions, evaluate form correctness through a reconstruction-error anomaly-detection criterion. A Retrieval-Augmented Generation (RAG) component built on the Gemma-2B-IT language model provides natural-language explanations and contextual coaching. Evaluation on a custom dataset of 9,414 videos covering three dumbbell exercises (bicep curl, shoulder press, lateral side raise) yields 95.9% classification accuracy, with a Bi-LSTM + Attention variant reaching 96.8%. The encoder-decoder reconstruction error separates correct and incorrect forms by an order of magnitude (mean ratio 10×–15×). The Chatbot achieves a perplexity of 6.8 and a mean human-rated helpfulness of 4.4/5 across 15 evaluators. The proposed multimodal architecture demonstrates a practical pathway toward intelligent home-fitness coaching.

Journal of Health Population and Nutrition
Carl von Ossietzky Universität Oldenburg (DE), Vellore Institute of Technology University (IN)
Openalex Percentile: Top 15%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.