Systematic benchmarking of evaluation paradigms, safety boundaries, and clinical reasoning gaps for multimodal large language models in rehabilitation

Abstract Motor rehabilitation and clinical biomechanical assessment rely on expert judgement to evaluate movement quality, pathological compensation, and spatiotemporal motor dynamics. Existing automated approaches mainly depend on conventional computer vision methods, such as skeletal keypoint-based analysis, but they provide limited clinical semantic interpretation and mechanistic reasoning. With the emergence of multimodal large language models (MLLMs), movement assessment may shift from geometric pose tracking toward end-to-end visual clinical reasoning. However, the ability of current MLLMs to identify fine-grained kinematic abnormalities and suppress visually unsupported hallucinations in rehabilitation videos remains insufficiently evaluated. This study aimed to develop a multidimensional video benchmark, ActionEval-MedBench , to systematically evaluate the clinical reasoning capability, safety boundaries, and failure modes of MLLMs in motor rehabilitation and clinical biomechanics. ActionEval-MedBench included 522 rigorously screened and de-identified videos, comprising 328 clinical upper-limb rehabilitation videos and 194 general functional biomechanics videos. Each video was paired with six multi-select multiple-choice questions, yielding 3132 video—MCQ assessment items. We designed a six-dimensional ActionEval protocol covering semantic action recognition, kinematic feature alignment, pathological compensation identification, spatiotemporal dynamics analysis, comprehensive movement quality and clinical decision-making, and visual grounding for anti-hallucination assessment. Expert consensus labels were established through a dual-track Delphi-style procedure. Twenty-six leading MLLMs were evaluated under standardized zero-shot prompts and structured response-parsing rules. The primary outcome was the cohort-weighted ActionEval composite score, while the D6 false-positive hallucination rate ( $$FP_{D6}$$ ) was used as an independent safety endpoint. Among 81,432 expected dimension-level response opportunities, 81,006 responses were successfully returned, whereas 426 were unavailable because of missing outputs or system-level failures. Among the returned responses, 80,931 were parsable and valid, while 75 were invalid or unparsable, yielding an overall valid-response rate of 99.38%. Gemini-3.1-Pro-Preview, Doubao-Seed-2.0-Lite-260428, Gemini-3.5-Flash, Claude-Opus-4-7, GLM-5V-Turbo, and GPT-5.5 formed a numerically high-performing cluster, with weighted ActionEval scores ranging from 66.17 to 65.46. The Friedman test showed significant overall performance differences among the 26 models ( $$\\chi ^2 = 2671.34$$ , $$P < 0.001$$ ), whereas Holm–Bonferroni-adjusted pairwise comparisons within the high-performing cluster were not statistically significant (all $$P > 0.05$$ ). Dimension-wise analysis revealed a marked semantic-to-clinical reasoning gap: models achieved strong performance in D1 semantic action recognition, with a mean F1-score of 91.1%, but declined substantially in higher-order clinical dimensions involving pathological compensation, spatiotemporal dynamics, and clinical decision-making, with mean F1-scores below 50% across D3–D5. Performance was lower in the clinical rehabilitation cohort than in the general functional biomechanics cohort, with an average drop of 9.44 percentage points; this difference was interpreted as a combined domain-shift burden between standardized functional videos and real-world fine-grained clinical rehabilitation videos rather than the isolated effect of pathology. In D6 negative-probe testing, nine models produced visually unsupported affirmative selections, with the highest $$FP_{D6}$$ reaching 6.51%. ActionEval-MedBench quantifies the cognitive boundaries and safety risks of current MLLMs in dynamic rehabilitation video understanding. Although MLLMs show promise for structured preliminary screening and clinical semantic reasoning, their limitations in detecting subtle pathological compensation and modeling spatiotemporal dynamics preclude their use as independent diagnostic tools. Future digital rehabilitation AI should follow a human-in-the-loop paradigm and incorporate domain-specific causal reasoning, kinematic-chain modeling, and rigorous visual-grounding safety probes.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-19
DOI
https://doi.org/10.1038/s41598-026-71999-w
Primary Topic
Stroke Rehabilitation and Recovery
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Systematic benchmarking of evaluation paradigms, safety boundaries, and clinical reasoning gaps for multimodal large language models in rehabilitation

Mengjian Qu, Yu Li, Xuan Zhang, Tao Zhu et al.
Scientific Reports
Stroke Rehabilitation and Recovery
article

Systematic benchmarking of evaluation paradigms, safety boundaries, and clinical reasoning gaps for multimodal large language models in rehabilitation

Mengjian Qu, Yu Li, Xuan Zhang, Tao Zhu, Jun Zhou, Ping Ye, Jing Liu
article en

Abstract

Abstract Motor rehabilitation and clinical biomechanical assessment rely on expert judgement to evaluate movement quality, pathological compensation, and spatiotemporal motor dynamics. Existing automated approaches mainly depend on conventional computer vision methods, such as skeletal keypoint-based analysis, but they provide limited clinical semantic interpretation and mechanistic reasoning. With the emergence of multimodal large language models (MLLMs), movement assessment may shift from geometric pose tracking toward end-to-end visual clinical reasoning. However, the ability of current MLLMs to identify fine-grained kinematic abnormalities and suppress visually unsupported hallucinations in rehabilitation videos remains insufficiently evaluated. This study aimed to develop a multidimensional video benchmark, ActionEval-MedBench , to systematically evaluate the clinical reasoning capability, safety boundaries, and failure modes of MLLMs in motor rehabilitation and clinical biomechanics. ActionEval-MedBench included 522 rigorously screened and de-identified videos, comprising 328 clinical upper-limb rehabilitation videos and 194 general functional biomechanics videos. Each video was paired with six multi-select multiple-choice questions, yielding 3132 video—MCQ assessment items. We designed a six-dimensional ActionEval protocol covering semantic action recognition, kinematic feature alignment, pathological compensation identification, spatiotemporal dynamics analysis, comprehensive movement quality and clinical decision-making, and visual grounding for anti-hallucination assessment. Expert consensus labels were established through a dual-track Delphi-style procedure. Twenty-six leading MLLMs were evaluated under standardized zero-shot prompts and structured response-parsing rules. The primary outcome was the cohort-weighted ActionEval composite score, while the D6 false-positive hallucination rate ( $$FP_{D6}$$ ) was used as an independent safety endpoint. Among 81,432 expected dimension-level response opportunities, 81,006 responses were successfully returned, whereas 426 were unavailable because of missing outputs or system-level failures. Among the returned responses, 80,931 were parsable and valid, while 75 were invalid or unparsable, yielding an overall valid-response rate of 99.38%. Gemini-3.1-Pro-Preview, Doubao-Seed-2.0-Lite-260428, Gemini-3.5-Flash, Claude-Opus-4-7, GLM-5V-Turbo, and GPT-5.5 formed a numerically high-performing cluster, with weighted ActionEval scores ranging from 66.17 to 65.46. The Friedman test showed significant overall performance differences among the 26 models ( $$\chi ^2 = 2671.34$$ , $$P < 0.001$$ ), whereas Holm–Bonferroni-adjusted pairwise comparisons within the high-performing cluster were not statistically significant (all $$P > 0.05$$ ). Dimension-wise analysis revealed a marked semantic-to-clinical reasoning gap: models achieved strong performance in D1 semantic action recognition, with a mean F1-score of 91.1%, but declined substantially in higher-order clinical dimensions involving pathological compensation, spatiotemporal dynamics, and clinical decision-making, with mean F1-scores below 50% across D3–D5. Performance was lower in the clinical rehabilitation cohort than in the general functional biomechanics cohort, with an average drop of 9.44 percentage points; this difference was interpreted as a combined domain-shift burden between standardized functional videos and real-world fine-grained clinical rehabilitation videos rather than the isolated effect of pathology. In D6 negative-probe testing, nine models produced visually unsupported affirmative selections, with the highest $$FP_{D6}$$ reaching 6.51%. ActionEval-MedBench quantifies the cognitive boundaries and safety risks of current MLLMs in dynamic rehabilitation video understanding. Although MLLMs show promise for structured preliminary screening and clinical semantic reasoning, their limitations in detecting subtle pathological compensation and modeling spatiotemporal dynamics preclude their use as independent diagnostic tools. Future digital rehabilitation AI should follow a human-in-the-loop paradigm and incorporate domain-specific causal reasoning, kinematic-chain modeling, and rigorous visual-grounding safety probes.

Scientific Reports
Hong Kong Polytechnic University (HK), First Affiliated Hospital of University of South China (CN), University of South China (CN)
Peace, Justice and strong institutions
Openalex Percentile: Top 14%
Stroke Rehabilitation and Recovery
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.