When Ground Truth Is the Bottleneck: How Self-Report Label Properties Constrain Multimodal User-State Estimation in Human–Robot Collaboration

Self-report instruments such as the NASA-TLX, STAI, and SAM routinely serve as ground-truth labels for multimodal user-state estimation, yet their measurement properties are rarely examined as constraints on prediction. In the MultiPhysio-HRC dataset (52 participants, 10 task conditions, 11 self-report variables), a generalizability study decomposed each label's variance into person, task, person-by-task, and error components, followed by leave-one-subject-out prediction. Task-level variance ranged from 9.7% to 62.9%. Treated as a dependability coefficient for the task-induced state, this proportion closely ordered cross-subject prediction (ρₛ = .900 for a physiological model; .982 with task identity added). Physiology improved all 11 labels (mean Δρ = +0.081). Labels fail for distinguishable reasons: person-dominated labels stay weak despite high repeatability, error-dominated labels are unstable at a single administration, and interaction-dominated labels may need person-specific calibration. The resulting limit is termed a measurement-and-construct-fit ceiling: a design-level constraint, not a bound on all predictors.

Authors

Institutions

Publication Details

Journal
International Journal of Human-Computer Interaction
Published
2026-09-17
DOI
https://doi.org/10.1080/10447318.2026.2732427
Primary Topic
Human-Automation Interaction and Safety
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

When Ground Truth Is the Bottleneck: How Self-Report Label Properties Constrain Multimodal User-State Estimation in Human–Robot Collaboration

Wonjoon Kim
International Journal of Human-Computer Interaction
Human-Automation Interaction and Safety
article

When Ground Truth Is the Bottleneck: How Self-Report Label Properties Constrain Multimodal User-State Estimation in Human–Robot Collaboration

Wonjoon Kim
article en

Abstract

Self-report instruments such as the NASA-TLX, STAI, and SAM routinely serve as ground-truth labels for multimodal user-state estimation, yet their measurement properties are rarely examined as constraints on prediction. In the MultiPhysio-HRC dataset (52 participants, 10 task conditions, 11 self-report variables), a generalizability study decomposed each label's variance into person, task, person-by-task, and error components, followed by leave-one-subject-out prediction. Task-level variance ranged from 9.7% to 62.9%. Treated as a dependability coefficient for the task-induced state, this proportion closely ordered cross-subject prediction (ρₛ = .900 for a physiological model; .982 with task identity added). Physiology improved all 11 labels (mean Δρ = +0.081). Labels fail for distinguishable reasons: person-dominated labels stay weak despite high repeatability, error-dominated labels are unstable at a single administration, and interaction-dominated labels may need person-specific calibration. The resulting limit is termed a measurement-and-construct-fit ceiling: a design-level constraint, not a bound on all predictors.

International Journal of Human-Computer Interaction
Dongduk Women's University (KR)
National Research Foundation of Korea
Openalex Percentile: Top 7%
Human-Automation Interaction and Safety
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

When Ground Truth Is the Bottleneck: How Self-Report Label Properties Constrain Multimodal User-State Estimation in Human–Robot Collaboration — Wonjoon Kim · International Journal of Human-Computer Interaction (2026) | TGRS Research Map | TGRS