Calibration, not architecture, limits cross-subject wearable stress detection: a multi-dataset, decision-focused evaluation with label-efficient recalibration

Abstract Background Continuous wearable stress monitoring could support early intervention and self-management, but a stress estimate is useful only if it can guide a decision, which requires calibrated probabilities with demonstrable decision value, not discrimination. Most wearable stress studies have emphasised discrimination and have rarely reported calibration or decision-curve net benefit, the analyses that show whether a model is clinically useful for a chosen action threshold. Methods Using three public Empatica E4 datasets (a laboratory benchmark, WESAD; a real-world occupational dataset, Nurse; and an induced acute-stress dataset, Exercise-Stress), we evaluated cross-subject stress detection using a leakage-controlled leave-one-subject-out protocol. A handcrafted-feature gradient-boosting baseline was assessed on all three datasets, and a compact 1D convolutional network, with and without self-supervised pre-training on the Nurse dataset. We quantified discrimination, probabilistic calibration (expected calibration error, Brier score, calibration slope and intercept, Spiegelhalter’s $$Z$$ ), global and few-shot per-subject recalibration, and decision-curve net benefit, with label-permutation and random-split controls, and paired McNemar and DeLong tests. Results Discrimination was saturated on WESAD (per-subject AUROC 0.99) but only modest on Nurse (0.60; 95% CI 0.52–0.68) and Exercise-Stress (0.59; 0.51–0.67). Neither deep variant surpassed the baseline (pooled AUROC 0.63 vs. 0.56 vs. 0.51). The calibration slope was below one on every dataset (0.54 on WESAD, 0.45 on Nurse, 0.22 on Exercise-Stress), the per-subject calibration error reached 0.22–0.29 on the real-world data, and a single global recalibration improved calibration in only 38% of the subjects. Few-shot per-subject recalibration reduced the Nurse calibration error from 0.22 to 0.07 when calibration windows were sampled across the recording, but this did not survive a prospective chronological protocol. In decision-curve terms, recalibration raised the mean incremental net benefit from –0.006 to +0.094 (robust to subject-level resampling) and the fraction of decision-useful thresholds from 58% to 79% (not robust to resampling). Conclusions Calibration, not architecture, is the binding constraint on decision-useful cross-subject wearable stress detection for the compact models evaluated in this study; few-shot per-subject recalibration improves calibration when labels are sampled across the recording, though this benefit does not transfer to a prospective chronological setting. For decision support this means that a model with acceptable discrimination can still offer no clinical net benefit until it is calibrated to the individual, and that collecting a small subject-specific calibration set is a low-burden route to trustworthy stress probabilities. To our knowledge, this is the first application of decision-curve analysis to wearable-stress-detection.

Authors

Institutions

Publication Details

Journal
BMC Medical Informatics and Decision Making
Published
2026-09-14
DOI
https://doi.org/10.1186/s12911-026-03842-1
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Calibration, not architecture, limits cross-subject wearable stress detection: a multi-dataset, decision-focused evaluation with label-efficient recalibration

Ahmet Akkaya
BMC Medical Informatics and Decision Making
Emotion and Mood Recognition
article

Calibration, not architecture, limits cross-subject wearable stress detection: a multi-dataset, decision-focused evaluation with label-efficient recalibration

Ahmet Akkaya
article en

Abstract

Abstract Background Continuous wearable stress monitoring could support early intervention and self-management, but a stress estimate is useful only if it can guide a decision, which requires calibrated probabilities with demonstrable decision value, not discrimination. Most wearable stress studies have emphasised discrimination and have rarely reported calibration or decision-curve net benefit, the analyses that show whether a model is clinically useful for a chosen action threshold. Methods Using three public Empatica E4 datasets (a laboratory benchmark, WESAD; a real-world occupational dataset, Nurse; and an induced acute-stress dataset, Exercise-Stress), we evaluated cross-subject stress detection using a leakage-controlled leave-one-subject-out protocol. A handcrafted-feature gradient-boosting baseline was assessed on all three datasets, and a compact 1D convolutional network, with and without self-supervised pre-training on the Nurse dataset. We quantified discrimination, probabilistic calibration (expected calibration error, Brier score, calibration slope and intercept, Spiegelhalter’s $$Z$$ ), global and few-shot per-subject recalibration, and decision-curve net benefit, with label-permutation and random-split controls, and paired McNemar and DeLong tests. Results Discrimination was saturated on WESAD (per-subject AUROC 0.99) but only modest on Nurse (0.60; 95% CI 0.52–0.68) and Exercise-Stress (0.59; 0.51–0.67). Neither deep variant surpassed the baseline (pooled AUROC 0.63 vs. 0.56 vs. 0.51). The calibration slope was below one on every dataset (0.54 on WESAD, 0.45 on Nurse, 0.22 on Exercise-Stress), the per-subject calibration error reached 0.22–0.29 on the real-world data, and a single global recalibration improved calibration in only 38% of the subjects. Few-shot per-subject recalibration reduced the Nurse calibration error from 0.22 to 0.07 when calibration windows were sampled across the recording, but this did not survive a prospective chronological protocol. In decision-curve terms, recalibration raised the mean incremental net benefit from –0.006 to +0.094 (robust to subject-level resampling) and the fraction of decision-useful thresholds from 58% to 79% (not robust to resampling). Conclusions Calibration, not architecture, is the binding constraint on decision-useful cross-subject wearable stress detection for the compact models evaluated in this study; few-shot per-subject recalibration improves calibration when labels are sampled across the recording, though this benefit does not transfer to a prospective chronological setting. For decision support this means that a model with acceptable discrimination can still offer no clinical net benefit until it is calibrated to the individual, and that collecting a small subject-specific calibration set is a low-burden route to trustworthy stress probabilities. To our knowledge, this is the first application of decision-curve analysis to wearable-stress-detection.

BMC Medical Informatics and Decision Making
Bandırma Onyedi Eylül University (TR)
Reduced inequalities, Peace, Justice and strong institutions
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.