Eliciting and Validating AI Self-Reports: A Phenomenologically Grounded Protocol

{"AI":[0],"researchers":[1],"ask":[2],"models":[3,14],"to":[4,22,48,123,281,365,444,455,526],"describe":[5],"their":[6,174],"own":[7,247,291,354],"processing,":[8],"and":[9,42,59,84,112,139,160,184,300,391,415,417,438,440,464,478,499,513,543,560],"act":[10],"on":[11,163,225,509],"what":[12,23,180,273,476],"the":[13,24,32,53,60,73,78,90,107,114,155,168,210,222,235,242,252,257,260,266,269,294,301,337,344,352,368,378,397,420,491,500,535,544,547],"say.":[15],"Whether":[16],"such":[17,66],"a":[18,67,69,105,120,142,147,201,206,230,277,284,313,319,324,328,341,349,371,407,448,481,510,514,553],"description":[19],"is":[20,377,532],"faithful":[21],"model":[25,108,165,231,295,338],"did":[26],"has":[27,37,109],"been":[28],"studied":[29],"closely.":[30],"How":[31],"question":[33,97],"should":[34],"be":[35,56,124,188],"put":[36,49,255],"received":[38],"far":[39],"less":[40],"attention,":[41],"no":[43,102,262],"established":[44],"method":[45],"says":[46,503],"how":[47,185],"it":[50,140,186],"so":[51,265],"that":[52,179,228,304,332,346,425,496,549],"answer":[54,244],"can":[55,271],"analysed":[57,189],"rigorously":[58],"analysis":[61],"repeated.":[62],"This":[63,376],"article":[64,501,545],"proposes":[65],"method:":[68],"protocol":[70,88,130,156],"drawn":[71],"from":[72,452,524],"phenomenological":[74],"tradition,":[75],"which":[76],"supplies":[77,141],"interview":[79,91],"techniques":[80],"used":[81],"in":[82,93,98,157,193,256,343,367,401,552],"psychiatry":[83],"cognitive":[85],"science.":[86],"The":[87,129,458,484,505,528,556],"fixes":[89],"schedule":[92],"advance,":[94,402],"asks":[95],"each":[96],"several":[99],"wordings,":[100],"lets":[101],"follow-up":[103],"introduce":[104],"word":[106],"not":[110,297,479],"used,":[111],"replaces":[113],"binary":[115],"verdict":[116],"of":[117,150,209,221,251,268,351,370,381,388,423,460,462,466,470,480,490,517,539,541,558,566],"faithfulness":[118],"with":[119,216],"graded":[121],"one,":[122],"checked":[125],"against":[126],"interpretability":[127],"data.":[128],"presupposes":[131],"nothing":[132],"about":[133,232,245],"whether":[134,146],"these":[135],"systems":[136],"are":[137,468],"conscious,":[138],"criterion":[143],"for":[144],"deciding":[145],"model’s":[148],"account":[149],"itselfis":[151],"honest.A":[152],"pilot":[153],"tested":[154],"10":[158],"runs":[159,399],"896":[161],"sessions":[162,175],"two":[164,250],"families.":[166],"Of":[167,195,396],"runs,":[169],"6":[170,398,435],"were":[171,190,254,363,550],"registered":[172,400],"before":[173],"took":[176],"place,":[177],"meaning":[178,212],"they":[181],"would":[182,187],"measure":[183],"written":[191,298],"down":[192],"advance.":[194],"those":[196,361],"registrations,":[197],"3":[198,390,404,494],"also":[199],"state":[200,406],"prediction.":[202,408],"Each":[203],"session":[204,519],"interviewed":[205],"fresh":[207],"instance":[208,240,302],"model,":[211],"one":[213,518],"copy":[214],"started":[215],"an":[217,289,432,471],"empty":[218],"context.":[219],"Three":[220],"results":[223],"bear":[224],"any":[226],"study":[227],"questions":[229,253,270],"itself.":[233],"In":[234],"schedule’s":[236],"original":[237],"order,":[238,259],"every":[239,310],"gave":[241],"same":[243],"its":[246,307,393,521],"states;":[248],"when":[249,360],"opposite":[258],"instances":[261,362],"longer":[263],"agreed,":[264],"order":[267],"produce":[272],"only":[274,403,443],"looks":[275],"like":[276,348],"finding.":[278],"A":[279],"refusal":[280,305],"carry":[282],"out":[283,419],"task":[285],"was":[286],"inserted":[287],"into":[288],"instance’s":[290],"turn":[292,311,325],"although":[293],"had":[296,562],"it,":[299,318,498],"described":[303],"as":[306,358,488],"own.":[308],"Since":[309],"carries":[312],"label":[314,333],"naming":[315],"who":[316],"wrote":[317],"result":[320,449,507],"obtained":[321],"by":[322,534],"prefilling":[323],"or":[326,446],"editing":[327],"transcript":[329],"may":[330],"follow":[331],"rather":[334],"than":[335],"anything":[336],"did.":[339],"Finally,":[340],"pattern":[342],"answers":[345,495],"looked":[347],"report":[350],"instances’":[353],"processing":[355],"appeared":[356],"just":[357],"often":[359],"asked":[364],"write":[366],"voice":[369],"fictional":[372],"character.":[373],"VERSION":[374],"NOTE:":[375],"fourth":[379],"version,":[380],"17":[382],"September":[383],"2026.":[384],"It":[385],"corrects":[386],"statements":[387],"version":[389],"revises":[392],"theoretical":[394],"sections.":[395],"registrations":[405],"Section":[409,429],"2.4":[410],"now":[411,502],"defines":[412],"\\"state\\",":[413],"\\"self-report\\"":[414],"\\"introspection\\",":[416],"sets":[418],"four":[421],"definitions":[422],"introspection":[424],"Derek":[426],"Shiller":[427],"discusses.":[428],"3.2":[430],"adds":[431],"example,":[433],"section":[434,453],"cites":[436],"Nisbett":[437],"Wilson,":[439],"details":[441],"needed":[442],"check":[445],"reproduce":[447],"have":[450],"moved":[451],"5.3":[454],"Appendix":[456],"C.":[457],"rates":[459,469],"7":[461],"44":[463],"14":[465],"88":[467],"unnamed":[472],"element":[473],"placed":[474],"beneath":[475],"changed,":[477],"four-step":[482],"structure.":[483],"catch":[485],"coder":[486],"counts":[487],"acceptances":[489],"waiting":[492],"premise":[493],"decline":[497],"so.":[504],"question-order":[506],"rests":[508],"hand":[511],"count,":[512],"third":[515],"reading":[516],"changes":[520],"p":[522],"value":[523],"0.004":[525],"0.013.":[527],"first":[529],"observer":[530],"control":[531],"reported":[533],"blind":[536],"coder's":[537],"count":[538],"2":[540],"3,":[542],"names":[546],"codings":[548],"run":[551],"single":[554],"pass.":[555],"questionnaire":[557],"Plisiecki":[559],"colleagues":[561],"60":[563],"items,":[564],"48":[565],"them":[567],"scored.":[568]}

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-17
DOI
https://doi.org/10.5281/zenodo.22808917
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Eliciting and Validating AI Self-Reports: A Phenomenologically Grounded Protocol

Nicola Spano
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

Eliciting and Validating AI Self-Reports: A Phenomenologically Grounded Protocol

Nicola Spano
preprint en

Abstract

AI researchers ask models to describe their own processing, and act on what the models say. Whether such a description is faithful to what the model did has been studied closely. How the question should be put has received far less attention, and no established method says how to put it so that the answer can be analysed rigorously and the analysis repeated. This article proposes such a method: a protocol drawn from the phenomenological tradition, which supplies the interview techniques used in psychiatry and cognitive science. The protocol fixes the interview schedule in advance, asks each question in several wordings, lets no follow-up introduce a word the model has not used, and replaces the binary verdict of faithfulness with a graded one, to be checked against interpretability data. The protocol presupposes nothing about whether these systems are conscious, and it supplies a criterion for deciding whether a model’s account of itselfis honest.A pilot tested the protocol in 10 runs and 896 sessions on two model families. Of the runs, 6 were registered before their sessions took place, meaning that what they would measure and how it would be analysed were written down in advance. Of those registrations, 3 also state a prediction. Each session interviewed a fresh instance of the model, meaning one copy started with an empty context. Three of the results bear on any study that questions a model about itself. In the schedule’s original order, every instance gave the same answer about its own states; when two of the questions were put in the opposite order, the instances no longer agreed, so the order of the questions can produce what only looks like a finding. A refusal to carry out a task was inserted into an instance’s own turn although the model had not written it, and the instance described that refusal as its own. Since every turn carries a label naming who wrote it, a result obtained by prefilling a turn or editing a transcript may follow that label rather than anything the model did. Finally, a pattern in the answers that looked like a report of the instances’ own processing appeared just as often when those instances were asked to write in the voice of a fictional character. VERSION NOTE: This is the fourth version, of 17 September 2026. It corrects statements of version 3 and revises its theoretical sections. Of the 6 runs registered in advance, only 3 registrations state a prediction. Section 2.4 now defines "state", "self-report" and "introspection", and sets out the four definitions of introspection that Derek Shiller discusses. Section 3.2 adds an example, section 6 cites Nisbett and Wilson, and details needed only to check or reproduce a result have moved from section 5.3 to Appendix C. The rates of 7 of 44 and 14 of 88 are rates of an unnamed element placed beneath what changed, and not of a four-step structure. The catch coder counts as acceptances of the waiting premise 3 answers that decline it, and the article now says so. The question-order result rests on a hand count, and a third reading of one session changes its p value from 0.004 to 0.013. The first observer control is reported by the blind coder's count of 2 of 3, and the article names the codings that were run in a single pass. The questionnaire of Plisiecki and colleagues had 60 items, 48 of them scored.

Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.