Emulation Diagnostics for Model Self-Reports: A Bridge Between Interpretability and Welfare Assessment

Two adjacent but methodologically separate substrates currently coexist inside Anthropic's published record. The first is a body of mechanistic-interpretability work, most prominently On the Biology of a Large Language Model [1], that produces attribution graphs over sparse activation features and cautions that verbal self-explanation can diverge from a recovered mechanism. The second is the §5 Claude Opus 4 welfare assessment in the May 2025 Opus 4 / Sonnet 4 system card [2], which discusses model self-reports about preferences, distress, and affect, including a reported "spiritual bliss" attractor state. Section 5.1 expressly limits form-to-semantics inference. This paper proposes four tests, including a propagation comparison between hedge-framed and hedge-stripped §5.2 prose plus factual and creative comparators. Each proposed test has a stated interpretation for null and non-null outcomes; none is reported as executed here.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.20368225
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Emulation Diagnostics for Model Self-Reports: A Bridge Between Interpretability and Welfare Assessment

Honeycutt, Edwin Marshall, III
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

Emulation Diagnostics for Model Self-Reports: A Bridge Between Interpretability and Welfare Assessment

Honeycutt, Edwin Marshall, III
preprint en

Abstract

Two adjacent but methodologically separate substrates currently coexist inside Anthropic's published record. The first is a body of mechanistic-interpretability work, most prominently On the Biology of a Large Language Model [1], that produces attribution graphs over sparse activation features and cautions that verbal self-explanation can diverge from a recovered mechanism. The second is the §5 Claude Opus 4 welfare assessment in the May 2025 Opus 4 / Sonnet 4 system card [2], which discusses model self-reports about preferences, distress, and affect, including a reported "spiritual bliss" attractor state. Section 5.1 expressly limits form-to-semantics inference. This paper proposes four tests, including a propagation comparison between hedge-framed and hedge-stripped §5.2 prose plus factual and creative comparators. Each proposed test has a stated interpretation for null and non-null outcomes; none is reported as executed here.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.