Emulation Diagnostics for Model Self-Reports: A Bridge Between Interpretability and Welfare Assessment
Two adjacent but methodologically separate substrates currently coexist inside Anthropic's published record. The first is a body of mechanistic-interpretability work, most prominently On the Biology of a Large Language Model [1], that produces attribution graphs over sparse activation features and cautions that verbal self-explanation can diverge from a recovered mechanism. The second is the §5 Claude Opus 4 welfare assessment in the May 2025 Opus 4 / Sonnet 4 system card [2], which discusses model self-reports about preferences, distress, and affect, including a reported "spiritual bliss" attractor state. Section 5.1 expressly limits form-to-semantics inference. This paper proposes four tests, including a propagation comparison between hedge-framed and hedge-stripped §5.2 prose plus factual and creative comparators. Each proposed test has a stated interpretation for null and non-null outcomes; none is reported as executed here.
Authors
- E. M. Honeycutt III
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23068607
- Primary Topic
- Embodied and Extended Cognition
- Type
- preprint