Machine Learning Models for Mental-Health Prediction from Actigraphy: A Systematic Appraisal and an Empirical Demonstration
Background and objective. The open actigraphy cohorts DEPRESJON (depression), PSYKOSE (schizophrenia) and HYPERAKTIV (ADHD), harmonised in 2025 as OBF-Psychiatric, are the most used public benchmarks for machine-learning (ML) detection of mental disorders from wearables. Reported accuracies have risen from the dataset authors' leave-one-patient-out 0.7 to above 0.9. We asked whether that reflects better models or weaker evaluation, and measured what the prevalent evaluation practices cost. Methods. Part - I appraises every published supervised model on the four cohorts with PROBAST+AI and TRIPOD+AI, extracting the unit of train/test splitting, external validation, calibration reporting, events per predictor, code availability, and whether the control group shared by DEPRESJON and PSYKOSE was double counted. Part~II re-evaluates the cohorts with models fixed a priori (logistic regression, random forest, XGBoost, 1D-CNN), ten seeds and subject-level bias-corrected bootstrap intervals, in six experiments: E1 record-wise versus subject-wise cross-validation on identical data; E2 frozen cross-cohort transfer with shared controls assigned to one cohort; E3 calibration (slope, intercept, Brier, ECE, reliability curves, recalibration); E4 sex strata; E5 transdiagnostic and five-class tasks on OBF-Psychiatric; E6 feature-group, resampling and cross-validation-scheme ablation. Results. Of 28 eligible published models, 7 split by participant, 9 by day or window and 12 did not state the unit; none reported calibration or validated on a second cohort; 4 released code. Record-wise splitting inflated window-level AUROC by 0.065 to 0.114 on DEPRESJON, 0.020 to 0.044 on PSYKOSE and 0.130 to 0.169 on HYPERAKTIV, where the honest estimate was at chance (0.44 to 0.45) and the leaky one reached 0.60 to 0.65 at participant level. Models trained on PSYKOSE lost 0.13 to 0.21 AUROC on DEPRESJON with calibration slopes of 0.27 to 0.63; honest slopes were 0.53 to 0.95 and ECE 0.06 to 0.22; recalibration from the training cohort did not repair the transfer. Participant-level events per predictor were 0.6 to 0.8. Honest performance was stable across feature subsets and cross-validation schemes (E6). Conclusions. Reported progress on these benchmarks is largely an artefact of evaluation design. We provide subject-wise, calibrated, cross-cohort baselines with public code, manifests and a TRIPOD+AI self-audit against which future claims on these cohorts can be checked.
Authors
- Roger-Nick Anaedevha (ORCID: https://orcid.org/0000-0003-3649-1561)
- Aminat A. Showole (ORCID: https://orcid.org/0000-0003-3387-6845)
Institutions
- University of Hafr Al-Batin (SA)
- University of Abuja (NG)
- National Research Nuclear University MEPhI (RU)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22772533
- Primary Topic
- Digital Mental Health Interventions
- Type
- article
- Field-Weighted Citation Impact
- 0.00