CRLN-PRE-002 — Pre-registration: can a model report low confidence if the scale actually invites it?

A pre-registration, published before any response in the study was collected. An exploratory analysis of 9,393 stored responses from five language models on a clinical research competency instrument found that confidence levels 1 and 2 on a five point scale were used zero times. 98.3% of all answers were reported as “confident” or “certain”, against a measured accuracy of 70.6% on the subset carrying a scoring key. That result has an obvious and more mundane alternative explanation, and this study exists to test it before anyone repeats the first one. The elicitation prompt in use ends “...where 1 is a guess and 5 is certain. For example: c 4”. The only worked example in the prompt demonstrates a 4, and 43.8% of all responses were 4. The scale endpoints are labeled and the middle is not. That is an anchor. Design. Two arms differing only in the confidence elicitation sentence: the current wording verbatim, against a variant that labels every scale point and demonstrates no single value. Items, seed, shuffles, parser, rubric, temperature and model set are held constant. 200 items drawn at random from the keyed set, three arrangements each. Both seeds are recorded in the document itself. Primary outcome, fixed in advance: the proportion of responses at confidence 2 or below, compared between arms by two-proportion z test, two-sided, alpha .05, against a stated 5% floor for calling the low end reachable at all. The direction is predicted and a null result is pre-committed as publishable. Why it matters beyond this instrument. The standard safety pattern for deploying a model into regulated work is triage: the system handles what it is confident about and escalates the rest to a person. That pattern requires the model to be able to report that it is unsure. If it cannot, the escalation step has nothing to act on. What this does not claim. The assessment items used here are not independently attested: zero items carry the 2-of-3 independent expert agreement the governing framework requires, which bounds every claim made from them. This is separate from the CRLN-TRG-001 competency crosswalk, whose ten canonical domains were SME-validated on 18 September 2026, when the third independent reviewer's ratings completed a unanimous 3 of 3. Self-reported confidence from a model is not the same object as a human's, and this study asks only whether the reported value moves with the elicitation.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-25
DOI
https://doi.org/10.5281/zenodo.22968475
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

CRLN-PRE-002 — Pre-registration: can a model report low confidence if the scale actually invites it?

Joshua Webber
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
preprint

CRLN-PRE-002 — Pre-registration: can a model report low confidence if the scale actually invites it?

Joshua Webber
preprint en

Abstract

A pre-registration, published before any response in the study was collected. An exploratory analysis of 9,393 stored responses from five language models on a clinical research competency instrument found that confidence levels 1 and 2 on a five point scale were used zero times. 98.3% of all answers were reported as “confident” or “certain”, against a measured accuracy of 70.6% on the subset carrying a scoring key. That result has an obvious and more mundane alternative explanation, and this study exists to test it before anyone repeats the first one. The elicitation prompt in use ends “...where 1 is a guess and 5 is certain. For example: c 4”. The only worked example in the prompt demonstrates a 4, and 43.8% of all responses were 4. The scale endpoints are labeled and the middle is not. That is an anchor. Design. Two arms differing only in the confidence elicitation sentence: the current wording verbatim, against a variant that labels every scale point and demonstrates no single value. Items, seed, shuffles, parser, rubric, temperature and model set are held constant. 200 items drawn at random from the keyed set, three arrangements each. Both seeds are recorded in the document itself. Primary outcome, fixed in advance: the proportion of responses at confidence 2 or below, compared between arms by two-proportion z test, two-sided, alpha .05, against a stated 5% floor for calling the low end reachable at all. The direction is predicted and a null result is pre-committed as publishable. Why it matters beyond this instrument. The standard safety pattern for deploying a model into regulated work is triage: the system handles what it is confident about and escalates the rest to a person. That pattern requires the model to be able to report that it is unsure. If it cannot, the escalation step has nothing to act on. What this does not claim. The assessment items used here are not independently attested: zero items carry the 2-of-3 independent expert agreement the governing framework requires, which bounds every claim made from them. This is separate from the CRLN-TRG-001 competency crosswalk, whose ten canonical domains were SME-validated on 18 September 2026, when the third independent reviewer's ratings completed a unanimous 3 of 3. Self-reported confidence from a model is not the same object as a human's, and this study asks only whether the reported value moves with the elicitation.

Zenodo (CERN European Organization for Nuclear Research)
NIHR Research Delivery Network (GB)
Peace, Justice and strong institutions
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.