When Ground Truth Is the Bottleneck: How Self-Report Label Properties Constrain Multimodal User-State Estimation in Human–Robot Collaboration
Self-report instruments such as the NASA-TLX, STAI, and SAM routinely serve as ground-truth labels for multimodal user-state estimation, yet their measurement properties are rarely examined as constraints on prediction. In the MultiPhysio-HRC dataset (52 participants, 10 task conditions, 11 self-report variables), a generalizability study decomposed each label's variance into person, task, person-by-task, and error components, followed by leave-one-subject-out prediction. Task-level variance ranged from 9.7% to 62.9%. Treated as a dependability coefficient for the task-induced state, this proportion closely ordered cross-subject prediction (ρₛ = .900 for a physiological model; .982 with task identity added). Physiology improved all 11 labels (mean Δρ = +0.081). Labels fail for distinguishable reasons: person-dominated labels stay weak despite high repeatability, error-dominated labels are unstable at a single administration, and interaction-dominated labels may need person-specific calibration. The resulting limit is termed a measurement-and-construct-fit ceiling: a design-level constraint, not a bound on all predictors.
Authors
- Wonjoon Kim (ORCID: https://orcid.org/0000-0001-5177-8072)
Institutions
- Dongduk Women's University (KR)
Publication Details
- Journal
- International Journal of Human-Computer Interaction
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1080/10447318.2026.2732427
- Primary Topic
- Human-Automation Interaction and Safety
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Research Foundation of Korea