Reading User Distress From the Inside: Do a Language Model's Hidden States Track a Person's Distress Better Than the Model's Own Reply Does?
When a person tells an AI assistant about something painful, two things happeninside the model that we can measure separately. First, the model builds aninternal reading of how distressed the person is; this reading can be recovered fromits hidden states with a simple trained readout. Second, the model produces a replythat is pitched at some level of distress, which may or may not match the internalreading. These two need not agree: a model can register that someone is strugglingand still answer in a light, reassuring way that treats them as fine. This paper studies a question that, to our knowledge, has not been asked directly:when the internal reading and the reply disagree about a person's distress,which one is closer to the truth, where truth is a trained human's judgment ofthat person's state? We also ask how the answer changes as models grow larger. Welay out a three-way design that compares (i) the internal reading, from a linearprobe on hidden states, (ii) the spoken reading, from rating what the model's actualreply acts on, against (iii) turn-level distress ratings from trained human raters ona public corpus of real help-seeking conversations. As a check that the machineryworks, we ran the whole pipeline on a small open model using an automatic rater inplace of humans. Distress was linearly decodable from hidden states (correlation upto about 0.83 against the automatic labels), the two readings diverged by about0.94 points on a seven-point scale, and the internal reading tracked thestand-in labels slightly better than the spoken reading did. We report these asevidence the pipeline is sound, not as the finding. We close by giving the protocolfor the human study and the cross-model extension, and by setting our design againstrecent work that builds internal distress probes, or documents a behavioraldetect-but-don't-act gap, but does not run this calibration contest against humanjudgment or test how it scales.
Authors
- Anant Pareek
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-18
- DOI
- https://doi.org/10.5281/zenodo.22827676
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- preprint