Federated learning for privacy-preserving collaborative health risk classification using wearable sensor data: validation study with elite football training load monitoring
Wearable sensor technologies are now widely adopted in professional sports for high-frequency health and performance monitoring during training and, in some cases, recovery periods, generating multimodal physiological and self-reported wellness data. These data carry high privacy sensitivity as protected personal health information, and in competitive sports they additionally hold tactical intelligence value. Consequently, multiorganization data sharing remains institutionally infeasible, creating data silos that constrain the scalability of data-driven health risk classification models. Federated learning (FL) has shown promise in health care for collaborative model training without centralizing raw data, but systematic empirical validation on athlete-level wearable health data with quantified, genuinely non–independently and identically distributed (non-IID) characteristics—spanning label, quantity, and covariate shift—remains absent. The present study therefore aimed to determine whether FL can achieve training load risk classification performance equivalent to centralized training while keeping raw athlete health data on-device, and to systematically investigate the key factors affecting FL efficacy in heterogeneous wearable health data scenarios. The publicly available SoccerMon dataset was used, comprising daily monitoring records from 50 elite female football players (all released players retained; filtering applied at the record level only) across 2 independent teams over the 2020 and 2021 seasons (36,550 raw records). A 21-dimensional multimodal feature matrix was constructed from 7 training load metrics derived from session rating of perceived exertion (sRPE) and its exponentially weighted moving averages (including the acute: chronic workload ratio [ACWR]), 7 subjective wellness indicators, and 7 binary missingness indicators. Using 28-day sliding windows, the future 7-day mean ACWR was classified into 3 risk levels (19,421 supervised samples). The federated clients exhibited three concurrent non-IID characteristics: label-distribution skew (SD of the high-risk class proportion across athlete clients 9.4% in TeamA and 21.7% in TeamB), quantity skew (mean 267–313 training windows per client), and between-team covariate shift that was statistically significant on all key training-load indicators but small in magnitude (Mann–Whitney r = .034–0.104), indicating that client-level rather than team-level heterogeneity dominates. Two-layer stacked long short-term memory (LSTM) networks served as client models within FedAvg and FedProx frameworks, compared against a centralized LSTM and 5 traditional machine learning baselines across single-team (K = 27), cross-team (K = 49), and team-level (K = 2) federated scenarios. All core experiments were repeated 30 times with Bonferroni-corrected paired t tests and Cohen d effect sizes. Team-level FedAvg retained 94.1% of centralized κ (0.401, SD 0.023 vs. 0.426, SD 0.018); cross-team athlete-level FedAvg retained 92.0% (κ = 0.392, SD 0.019); and single-team FedAvg retained 90.0% (κ = 0.366, SD 0.021). Two one-sided tests (TOST) established equivalence within a ± 0.05 κ practical bound between FL and centralized training in the cross-team and team-level scenarios (both TOST P ≤ .001), and between FedAvg and FedProx in all three scenarios (all TOST P < .001). High-risk recall remained within a narrow band across scenarios (FedAvg 0.767–0.828 vs. centralized LSTM 0.796–0.806). A naive persistence baseline reached κ of only 0.253–0.262, exceeded by FedAvg by 44.6%–53.1%, indicating that the model captures substantial structure beyond trivial autoregression. FedAvg and FedProx performed equivalently, with FedAvg approximately twice as computationally efficient. Load-only training features (7 dimensions) matched the full 21-dimensional set, and sRPE-derived internal load features contributed the majority of predictive importance. FL enables training load risk classification statistically equivalent to centralized training within a ± 0.05 κ practical bound, without sharing raw wearable health data, providing a viable privacy-preserving pathway for cross-organization sports analytics and health monitoring collaboration. These findings extend FL validation from clinical health care to athlete-level wearable health data with genuine non-IID characteristics, and remain robust under strictly causal data preprocessing.
Authors
- Bo Yang (ORCID: https://orcid.org/0000-0002-4446-2931)
- Xin Zhang
- Ziyu Liu
- Lianzhen Ma
- Xiangwu Li
- Zhao Gao
- Yupeng Shen
- Chijun Mu
- Zijun Liu
Institutions
- South China Normal University (CN)
- Guangzhou Sport University (CN)
- Art Innovation (Netherlands) (NL)
Publication Details
- Journal
- BMC Sports Science Medicine and Rehabilitation
- Published
- 2026-08-24
- DOI
- https://doi.org/10.1186/s13102-026-02020-0
- Primary Topic
- Physical Activity and Health
- Type
- article
- Field-Weighted Citation Impact
- 0.00