Machine Learning Models for Mental-Health Prediction from Actigraphy: A Systematic Appraisal and an Empirical Demonstration

Background and objective. The open actigraphy cohorts DEPRESJON (depression), PSYKOSE (schizophrenia) and HYPERAKTIV (ADHD), harmonised in 2025 as OBF-Psychiatric, are the most used public benchmarks for machine-learning (ML) detection of mental disorders from wearables. Reported accuracies have risen from the dataset authors' leave-one-patient-out 0.7 to above 0.9. We asked whether that reflects better models or weaker evaluation, and measured what the prevalent evaluation practices cost. Methods. Part - I appraises every published supervised model on the four cohorts with PROBAST+AI and TRIPOD+AI, extracting the unit of train/test splitting, external validation, calibration reporting, events per predictor, code availability, and whether the control group shared by DEPRESJON and PSYKOSE was double counted. Part~II re-evaluates the cohorts with models fixed a priori (logistic regression, random forest, XGBoost, 1D-CNN), ten seeds and subject-level bias-corrected bootstrap intervals, in six experiments: E1 record-wise versus subject-wise cross-validation on identical data; E2 frozen cross-cohort transfer with shared controls assigned to one cohort; E3 calibration (slope, intercept, Brier, ECE, reliability curves, recalibration); E4 sex strata; E5 transdiagnostic and five-class tasks on OBF-Psychiatric; E6 feature-group, resampling and cross-validation-scheme ablation. Results. Of 28 eligible published models, 7 split by participant, 9 by day or window and 12 did not state the unit; none reported calibration or validated on a second cohort; 4 released code. Record-wise splitting inflated window-level AUROC by 0.065 to 0.114 on DEPRESJON, 0.020 to 0.044 on PSYKOSE and 0.130 to 0.169 on HYPERAKTIV, where the honest estimate was at chance (0.44 to 0.45) and the leaky one reached 0.60 to 0.65 at participant level. Models trained on PSYKOSE lost 0.13 to 0.21 AUROC on DEPRESJON with calibration slopes of 0.27 to 0.63; honest slopes were 0.53 to 0.95 and ECE 0.06 to 0.22; recalibration from the training cohort did not repair the transfer. Participant-level events per predictor were 0.6 to 0.8. Honest performance was stable across feature subsets and cross-validation schemes (E6). Conclusions. Reported progress on these benchmarks is largely an artefact of evaluation design. We provide subject-wise, calibrated, cross-cohort baselines with public code, manifests and a TRIPOD+AI self-audit against which future claims on these cohorts can be checked.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-15
DOI
https://doi.org/10.5281/zenodo.22772533
Primary Topic
Digital Mental Health Interventions
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Machine Learning Models for Mental-Health Prediction from Actigraphy: A Systematic Appraisal and an Empirical Demonstration

Roger-Nick Anaedevha, Aminat A. Showole
Zenodo (CERN European Organization for Nuclear Research)
Digital Mental Health Interventions
article

Machine Learning Models for Mental-Health Prediction from Actigraphy: A Systematic Appraisal and an Empirical Demonstration

Roger-Nick Anaedevha, Aminat A. Showole
article en

Abstract

Background and objective. The open actigraphy cohorts DEPRESJON (depression), PSYKOSE (schizophrenia) and HYPERAKTIV (ADHD), harmonised in 2025 as OBF-Psychiatric, are the most used public benchmarks for machine-learning (ML) detection of mental disorders from wearables. Reported accuracies have risen from the dataset authors' leave-one-patient-out 0.7 to above 0.9. We asked whether that reflects better models or weaker evaluation, and measured what the prevalent evaluation practices cost. Methods. Part - I appraises every published supervised model on the four cohorts with PROBAST+AI and TRIPOD+AI, extracting the unit of train/test splitting, external validation, calibration reporting, events per predictor, code availability, and whether the control group shared by DEPRESJON and PSYKOSE was double counted. Part~II re-evaluates the cohorts with models fixed a priori (logistic regression, random forest, XGBoost, 1D-CNN), ten seeds and subject-level bias-corrected bootstrap intervals, in six experiments: E1 record-wise versus subject-wise cross-validation on identical data; E2 frozen cross-cohort transfer with shared controls assigned to one cohort; E3 calibration (slope, intercept, Brier, ECE, reliability curves, recalibration); E4 sex strata; E5 transdiagnostic and five-class tasks on OBF-Psychiatric; E6 feature-group, resampling and cross-validation-scheme ablation. Results. Of 28 eligible published models, 7 split by participant, 9 by day or window and 12 did not state the unit; none reported calibration or validated on a second cohort; 4 released code. Record-wise splitting inflated window-level AUROC by 0.065 to 0.114 on DEPRESJON, 0.020 to 0.044 on PSYKOSE and 0.130 to 0.169 on HYPERAKTIV, where the honest estimate was at chance (0.44 to 0.45) and the leaky one reached 0.60 to 0.65 at participant level. Models trained on PSYKOSE lost 0.13 to 0.21 AUROC on DEPRESJON with calibration slopes of 0.27 to 0.63; honest slopes were 0.53 to 0.95 and ECE 0.06 to 0.22; recalibration from the training cohort did not repair the transfer. Participant-level events per predictor were 0.6 to 0.8. Honest performance was stable across feature subsets and cross-validation schemes (E6). Conclusions. Reported progress on these benchmarks is largely an artefact of evaluation design. We provide subject-wise, calibrated, cross-cohort baselines with public code, manifests and a TRIPOD+AI self-audit against which future claims on these cohorts can be checked.

Zenodo (CERN European Organization for Nuclear Research)
University of Hafr Al-Batin (SA), University of Abuja (NG), National Research Nuclear University MEPhI (RU)
Openalex Percentile: Top 9%
Digital Mental Health Interventions
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.