Calibration, not architecture, limits cross-subject wearable stress detection: a multi-dataset, decision-focused evaluation with label-efficient recalibration
Abstract Background Continuous wearable stress monitoring could support early intervention and self-management, but a stress estimate is useful only if it can guide a decision, which requires calibrated probabilities with demonstrable decision value, not discrimination. Most wearable stress studies have emphasised discrimination and have rarely reported calibration or decision-curve net benefit, the analyses that show whether a model is clinically useful for a chosen action threshold. Methods Using three public Empatica E4 datasets (a laboratory benchmark, WESAD; a real-world occupational dataset, Nurse; and an induced acute-stress dataset, Exercise-Stress), we evaluated cross-subject stress detection using a leakage-controlled leave-one-subject-out protocol. A handcrafted-feature gradient-boosting baseline was assessed on all three datasets, and a compact 1D convolutional network, with and without self-supervised pre-training on the Nurse dataset. We quantified discrimination, probabilistic calibration (expected calibration error, Brier score, calibration slope and intercept, Spiegelhalter’s $$Z$$ ), global and few-shot per-subject recalibration, and decision-curve net benefit, with label-permutation and random-split controls, and paired McNemar and DeLong tests. Results Discrimination was saturated on WESAD (per-subject AUROC 0.99) but only modest on Nurse (0.60; 95% CI 0.52–0.68) and Exercise-Stress (0.59; 0.51–0.67). Neither deep variant surpassed the baseline (pooled AUROC 0.63 vs. 0.56 vs. 0.51). The calibration slope was below one on every dataset (0.54 on WESAD, 0.45 on Nurse, 0.22 on Exercise-Stress), the per-subject calibration error reached 0.22–0.29 on the real-world data, and a single global recalibration improved calibration in only 38% of the subjects. Few-shot per-subject recalibration reduced the Nurse calibration error from 0.22 to 0.07 when calibration windows were sampled across the recording, but this did not survive a prospective chronological protocol. In decision-curve terms, recalibration raised the mean incremental net benefit from –0.006 to +0.094 (robust to subject-level resampling) and the fraction of decision-useful thresholds from 58% to 79% (not robust to resampling). Conclusions Calibration, not architecture, is the binding constraint on decision-useful cross-subject wearable stress detection for the compact models evaluated in this study; few-shot per-subject recalibration improves calibration when labels are sampled across the recording, though this benefit does not transfer to a prospective chronological setting. For decision support this means that a model with acceptable discrimination can still offer no clinical net benefit until it is calibrated to the individual, and that collecting a small subject-specific calibration set is a low-burden route to trustworthy stress probabilities. To our knowledge, this is the first application of decision-curve analysis to wearable-stress-detection.
Authors
- Ahmet Akkaya (ORCID: https://orcid.org/0000-0003-4836-2310)
Institutions
- Bandırma Onyedi Eylül University (TR)
Publication Details
- Journal
- BMC Medical Informatics and Decision Making
- Published
- 2026-09-14
- DOI
- https://doi.org/10.1186/s12911-026-03842-1
- Primary Topic
- Emotion and Mood Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00