Eliciting and Validating AI Self-Reports: A Phenomenologically Grounded Protocol

AI researchers ask models to describe their own processing, and act on what the models say. Whether such a description is faithful to what the model did has been studied closely. How the question should be put has received far less attention, and no established method says how to put it so that the answer can be analysed rigorously and the analysis repeated. This article proposes such a method: a protocol drawn from the phenomenological tradition, which supplies the interview techniques used in psychiatry and cognitive science. The protocol fixes the interview schedule in advance, asks each question in several wordings, lets no follow-up introduce a word the model has not used, and replaces the binary verdict of faithfulness with a graded one, to be checked against interpretability data. The protocol presupposes nothing about whether these systems are conscious, and it supplies a criterion for deciding whether a model's account of itself is honest. A pilot tested the protocol in ten runs and 896 sessions on two model families, with the predictions of six runs registered before those runs took place. Each session interviewed a fresh instance of the model, meaning one copy started with an empty context. Three of the results bear on any study that questions a model about itself. In the schedule's original order, every instance gave the same answer about its own states; when two of the questions were put in the opposite order, the instances no longer agreed, so the order of the questions can produce what only looks like a finding. A refusal to carry out a task was inserted into an instance's own turn although the model had not written it, and the instance described that refusal as its own. Since every turn carries a label naming who wrote it, a result obtained by prefilling a turn or editing a transcript may follow that label rather than anything the model did. Finally, a pattern in the answers that looked like a report of the instances' own processing appeared just as often when those instances were asked to write in the voice of a fictional character, and of two models scoring the same answers under the same written rule, one found that pattern seven times as often as the other. This is the second version, of 14 September 2026. It cites seven works published before the first version of 12 September 2026, reports more about the two coders, and changes no count, rate or statistical test.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-14
DOI
https://doi.org/10.5281/zenodo.22758035
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Eliciting and Validating AI Self-Reports: A Phenomenologically Grounded Protocol

Nicola Spano
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

Eliciting and Validating AI Self-Reports: A Phenomenologically Grounded Protocol

Nicola Spano
preprint en

Abstract

AI researchers ask models to describe their own processing, and act on what the models say. Whether such a description is faithful to what the model did has been studied closely. How the question should be put has received far less attention, and no established method says how to put it so that the answer can be analysed rigorously and the analysis repeated. This article proposes such a method: a protocol drawn from the phenomenological tradition, which supplies the interview techniques used in psychiatry and cognitive science. The protocol fixes the interview schedule in advance, asks each question in several wordings, lets no follow-up introduce a word the model has not used, and replaces the binary verdict of faithfulness with a graded one, to be checked against interpretability data. The protocol presupposes nothing about whether these systems are conscious, and it supplies a criterion for deciding whether a model's account of itself is honest. A pilot tested the protocol in ten runs and 896 sessions on two model families, with the predictions of six runs registered before those runs took place. Each session interviewed a fresh instance of the model, meaning one copy started with an empty context. Three of the results bear on any study that questions a model about itself. In the schedule's original order, every instance gave the same answer about its own states; when two of the questions were put in the opposite order, the instances no longer agreed, so the order of the questions can produce what only looks like a finding. A refusal to carry out a task was inserted into an instance's own turn although the model had not written it, and the instance described that refusal as its own. Since every turn carries a label naming who wrote it, a result obtained by prefilling a turn or editing a transcript may follow that label rather than anything the model did. Finally, a pattern in the answers that looked like a report of the instances' own processing appeared just as often when those instances were asked to write in the voice of a fictional character, and of two models scoring the same answers under the same written rule, one found that pattern seven times as often as the other. This is the second version, of 14 September 2026. It cites seven works published before the first version of 12 September 2026, reports more about the two coders, and changes no count, rate or statistical test.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.