Predicting subsequent-year physical fitness assessment failure in university students using longitudinal data and explainable machine learning

Annual university physical fitness assessments are primarily used to describe students’ current fitness status, but their value for predicting failure in the subsequent year remains unclear. This study evaluated whether routinely collected physical fitness measurements could predict subsequent-year failure using a temporally separated modeling framework. This retrospective longitudinal prediction study used annual physical fitness assessment records from a university in China collected from 2023 to 2025. We developed models using 11,430 linked 2023–2024 student-year pairs and evaluated them on an out-of-time test dataset of 11,427 linked 2024–2025 student-year pairs. We defined subsequent-year failure as a total physical fitness score below 60. Predictors from the preceding year included body mass index, vital capacity, 50 m sprint time, standing long jump distance, sit-and-reach distance, endurance run time, and sex-specific muscular strength. We compared logistic regression, random forest, and XGBoost using discrimination, calibration, and classification metrics at a prespecified probability threshold of 0.50. We used SHapley Additive exPlanations to characterize feature contributions within the fitted XGBoost model. Subsequent-year failure occurred in 6.39% of the model-development dataset and 6.30% (720/11,427) of the out-of-time test dataset. Logistic regression showed the highest discrimination and recall (ROC-AUC = 0.8167; PR-AUC = 0.3227; recall = 0.7569). Random forest showed the highest accuracy (0.9320), F1 score (0.3573), and the lowest Brier score (0.0578), but its recall was 0.3000, and its calibration slope was 0.4985. XGBoost achieved an ROC-AUC of 0.7935 and a recall of 0.5458. In the XGBoost SHAP summary plot, BMI, endurance run time, standing long jump, and vital capacity showed the largest overall contributions to model output. Sensitivity analyses showed lower predictive performance under percentile-based definitions of subsequent-year failure. Although predictor measurements preceded the outcomes, the predictors and subsequent-year total fitness score were derived from the same physical fitness assessment framework; therefore, persistence of the underlying fitness construct and conceptual predictor–outcome overlap may have contributed to the observed performance. Routinely collected physical fitness measurements provided useful discrimination for subsequent-year failure, although model performance varied by evaluation metric and outcome definition. We conducted the out-of-time evaluation within a single university, and it should not be interpreted as independent external validation. External validation, local calibration, threshold assessment, and prospective evaluation are required before using the models for routine student monitoring or intervention.

Authors

Institutions

Publication Details

Journal
BMC Public Health
Published
2026-10-05
DOI
https://doi.org/10.1186/s12889-026-29742-7
Primary Topic
Physical Activity and Health
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Predicting subsequent-year physical fitness assessment failure in university students using longitudinal data and explainable machine learning

Lan Luo, Baixia Li, Fawei Huang
BMC Public Health
Physical Activity and Health
article

Predicting subsequent-year physical fitness assessment failure in university students using longitudinal data and explainable machine learning

Lan Luo, Baixia Li, Fawei Huang
article en

Abstract

Annual university physical fitness assessments are primarily used to describe students’ current fitness status, but their value for predicting failure in the subsequent year remains unclear. This study evaluated whether routinely collected physical fitness measurements could predict subsequent-year failure using a temporally separated modeling framework. This retrospective longitudinal prediction study used annual physical fitness assessment records from a university in China collected from 2023 to 2025. We developed models using 11,430 linked 2023–2024 student-year pairs and evaluated them on an out-of-time test dataset of 11,427 linked 2024–2025 student-year pairs. We defined subsequent-year failure as a total physical fitness score below 60. Predictors from the preceding year included body mass index, vital capacity, 50 m sprint time, standing long jump distance, sit-and-reach distance, endurance run time, and sex-specific muscular strength. We compared logistic regression, random forest, and XGBoost using discrimination, calibration, and classification metrics at a prespecified probability threshold of 0.50. We used SHapley Additive exPlanations to characterize feature contributions within the fitted XGBoost model. Subsequent-year failure occurred in 6.39% of the model-development dataset and 6.30% (720/11,427) of the out-of-time test dataset. Logistic regression showed the highest discrimination and recall (ROC-AUC = 0.8167; PR-AUC = 0.3227; recall = 0.7569). Random forest showed the highest accuracy (0.9320), F1 score (0.3573), and the lowest Brier score (0.0578), but its recall was 0.3000, and its calibration slope was 0.4985. XGBoost achieved an ROC-AUC of 0.7935 and a recall of 0.5458. In the XGBoost SHAP summary plot, BMI, endurance run time, standing long jump, and vital capacity showed the largest overall contributions to model output. Sensitivity analyses showed lower predictive performance under percentile-based definitions of subsequent-year failure. Although predictor measurements preceded the outcomes, the predictors and subsequent-year total fitness score were derived from the same physical fitness assessment framework; therefore, persistence of the underlying fitness construct and conceptual predictor–outcome overlap may have contributed to the observed performance. Routinely collected physical fitness measurements provided useful discrimination for subsequent-year failure, although model performance varied by evaluation metric and outcome definition. We conducted the out-of-time evaluation within a single university, and it should not be interpreted as independent external validation. External validation, local calibration, threshold assessment, and prospective evaluation are required before using the models for routine student monitoring or intervention.

BMC Public Health
East China University of Science and Technology (CN), Sichuan Technology and Business University (CN), International University (KH)
Good health and well-being
Openalex Percentile: Top 13%
Physical Activity and Health
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.