Estimating presenteeism from repeated smartphone-based multimodal behavioral responses

Presenteeism, attending work despite physical or mental health problems, is a major source of productivity loss worldwide, yet its early detection in everyday settings remains challenging. Conventional self-report instruments are time-consuming and require psychiatrist interpretation, which limits their scalability. We developed a multimodal machine-learning model that estimates work-function impairment as an indicator of presenteeism from brief smartphone-based dialogues. 38 employees recorded short self-report videos (10–30 s) using front-facing smartphone cameras, which were independently rated on an ordinal three-level work-function scale (Healthy / Moderate / Unwell) by three psychiatrists; the consensus label served as supervision. We tested whether integrating multimodal behavioural signals (acoustic, facial, and linguistic) would provide additional cues beyond linguistic content alone, particularly when verbal cues are sparse. Acoustic, facial, and linguistic features were extracted from each clip and integrated using a three-stream Attention-based Multiple Instance Learning (Attention-MIL) framework with learned late fusion, an ordinal-regression head, and class-prior logit adjustment. Generalisation was assessed under a participant-disjoint protocol (25-fold repeated GroupKFold; n = 1,768 out-of-fold clips from 29 participants). On the Full configuration, the proposed multimodal framework achieved macro-F1 = 0.772 (95% CI [0.697, 0.817]), accuracy = 0.840, macro-AUROC = 0.940, and low expected calibration error (ECE ≈ 0.024 ). At the aggregate level, the four text-containing configurations (Full / Audio+Text / Face+Text / Text-only) were mutually indistinguishable on macro-F1 (Holm–Bonferroni p adj > 0.40), whereas text-free configurations performed substantially worse ( Δ < − 0.33 , p adj < 0.001), confirming the necessity of the linguistic channel. In an exploratory subgroup analysis, Audio+Text outperformed Text-only in the clinically ambiguous subgroup—short-speech responses from non-Healthy participants ( n = 77; Δ macro-F1 =+0.030, exploratory, requires replication). These findings provide localised but clinically meaningful support for the multimodal hypothesis under low-information conditions and motivate prospective replication.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-16
DOI
https://doi.org/10.1371/journal.pone.0354342
Primary Topic
Digital Mental Health Interventions
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Estimating presenteeism from repeated smartphone-based multimodal behavioral responses

Kenji Suzuki, Masakazu Hirokawa, Shotaro Doki, Taiga Noguchi et al.
PLoS ONE
Digital Mental Health Interventions
article

Estimating presenteeism from repeated smartphone-based multimodal behavioral responses

Kenji Suzuki, Masakazu Hirokawa, Shotaro Doki, Taiga Noguchi, Soma Nishimura, Katsuya Hotta, Shota Matsumoto, Naoko Kouda, Yuya Iwata
article en

Abstract

Presenteeism, attending work despite physical or mental health problems, is a major source of productivity loss worldwide, yet its early detection in everyday settings remains challenging. Conventional self-report instruments are time-consuming and require psychiatrist interpretation, which limits their scalability. We developed a multimodal machine-learning model that estimates work-function impairment as an indicator of presenteeism from brief smartphone-based dialogues. 38 employees recorded short self-report videos (10–30 s) using front-facing smartphone cameras, which were independently rated on an ordinal three-level work-function scale (Healthy / Moderate / Unwell) by three psychiatrists; the consensus label served as supervision. We tested whether integrating multimodal behavioural signals (acoustic, facial, and linguistic) would provide additional cues beyond linguistic content alone, particularly when verbal cues are sparse. Acoustic, facial, and linguistic features were extracted from each clip and integrated using a three-stream Attention-based Multiple Instance Learning (Attention-MIL) framework with learned late fusion, an ordinal-regression head, and class-prior logit adjustment. Generalisation was assessed under a participant-disjoint protocol (25-fold repeated GroupKFold; n = 1,768 out-of-fold clips from 29 participants). On the Full configuration, the proposed multimodal framework achieved macro-F1 = 0.772 (95% CI [0.697, 0.817]), accuracy = 0.840, macro-AUROC = 0.940, and low expected calibration error (ECE ≈ 0.024 ). At the aggregate level, the four text-containing configurations (Full / Audio+Text / Face+Text / Text-only) were mutually indistinguishable on macro-F1 (Holm–Bonferroni p adj > 0.40), whereas text-free configurations performed substantially worse ( Δ < − 0.33 , p adj < 0.001), confirming the necessity of the linguistic channel. In an exploratory subgroup analysis, Audio+Text outperformed Text-only in the clinically ambiguous subgroup—short-speech responses from non-Healthy participants ( n = 77; Δ macro-F1 =+0.030, exploratory, requires replication). These findings provide localised but clinically meaningful support for the multimodal hypothesis under low-information conditions and motivate prospective replication.

PLoS ONEVol. 21(9)
University of Tsukuba (JP), Applied Minds (United States) (US), China Communications Construction Company (China) (CN)
Decent work and economic growth
Openalex Percentile: Top 9%
Digital Mental Health Interventions
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.