Reading User Distress From the Inside: Do a Language Model's Hidden States Track a Person's Distress Better Than the Model's Own Reply Does?

When a person tells an AI assistant about something painful, two things happeninside the model that we can measure separately. First, the model builds aninternal reading of how distressed the person is; this reading can be recovered fromits hidden states with a simple trained readout. Second, the model produces a replythat is pitched at some level of distress, which may or may not match the internalreading. These two need not agree: a model can register that someone is strugglingand still answer in a light, reassuring way that treats them as fine. This paper studies a question that, to our knowledge, has not been asked directly:when the internal reading and the reply disagree about a person's distress,which one is closer to the truth, where truth is a trained human's judgment ofthat person's state? We also ask how the answer changes as models grow larger. Welay out a three-way design that compares (i) the internal reading, from a linearprobe on hidden states, (ii) the spoken reading, from rating what the model's actualreply acts on, against (iii) turn-level distress ratings from trained human raters ona public corpus of real help-seeking conversations. As a check that the machineryworks, we ran the whole pipeline on a small open model using an automatic rater inplace of humans. Distress was linearly decodable from hidden states (correlation upto about 0.83 against the automatic labels), the two readings diverged by about0.94 points on a seven-point scale, and the internal reading tracked thestand-in labels slightly better than the spoken reading did. We report these asevidence the pipeline is sound, not as the finding. We close by giving the protocolfor the human study and the cross-model extension, and by setting our design againstrecent work that builds internal distress probes, or documents a behavioraldetect-but-don't-act gap, but does not run this calibration contest against humanjudgment or test how it scales.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-18
DOI
https://doi.org/10.5281/zenodo.22827676
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Reading User Distress From the Inside: Do a Language Model's Hidden States Track a Person's Distress Better Than the Model's Own Reply Does?

Anant Pareek
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

Reading User Distress From the Inside: Do a Language Model's Hidden States Track a Person's Distress Better Than the Model's Own Reply Does?

Anant Pareek
preprint en

Abstract

When a person tells an AI assistant about something painful, two things happeninside the model that we can measure separately. First, the model builds aninternal reading of how distressed the person is; this reading can be recovered fromits hidden states with a simple trained readout. Second, the model produces a replythat is pitched at some level of distress, which may or may not match the internalreading. These two need not agree: a model can register that someone is strugglingand still answer in a light, reassuring way that treats them as fine. This paper studies a question that, to our knowledge, has not been asked directly:when the internal reading and the reply disagree about a person's distress,which one is closer to the truth, where truth is a trained human's judgment ofthat person's state? We also ask how the answer changes as models grow larger. Welay out a three-way design that compares (i) the internal reading, from a linearprobe on hidden states, (ii) the spoken reading, from rating what the model's actualreply acts on, against (iii) turn-level distress ratings from trained human raters ona public corpus of real help-seeking conversations. As a check that the machineryworks, we ran the whole pipeline on a small open model using an automatic rater inplace of humans. Distress was linearly decodable from hidden states (correlation upto about 0.83 against the automatic labels), the two readings diverged by about0.94 points on a seven-point scale, and the internal reading tracked thestand-in labels slightly better than the spoken reading did. We report these asevidence the pipeline is sound, not as the finding. We close by giving the protocolfor the human study and the cross-model extension, and by setting our design againstrecent work that builds internal distress probes, or documents a behavioraldetect-but-don't-act gap, but does not run this calibration contest against humanjudgment or test how it scales.

Zenodo (CERN European Organization for Nuclear Research)
Quality Education
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Reading User Distress From the Inside: Do a Language Model's Hidden States Track a Person's Distress Better Than the Model's Own Reply Does? — Anant Pareek · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS