CRLN-PRE-002 — Pre-registration: can a model report low confidence if the scale actually invites it?
A pre-registration, published before any response in the study was collected. An exploratory analysis of 9,393 stored responses from five language models on a clinical research competency instrument found that confidence levels 1 and 2 on a five point scale were used zero times. 98.3% of all answers were reported as “confident” or “certain”, against a measured accuracy of 70.6% on the subset carrying a scoring key. That result has an obvious and more mundane alternative explanation, and this study exists to test it before anyone repeats the first one. The elicitation prompt in use ends “...where 1 is a guess and 5 is certain. For example: c 4”. The only worked example in the prompt demonstrates a 4, and 43.8% of all responses were 4. The scale endpoints are labeled and the middle is not. That is an anchor. Design. Two arms differing only in the confidence elicitation sentence: the current wording verbatim, against a variant that labels every scale point and demonstrates no single value. Items, seed, shuffles, parser, rubric, temperature and model set are held constant. 200 items drawn at random from the keyed set, three arrangements each. Both seeds are recorded in the document itself. Primary outcome, fixed in advance: the proportion of responses at confidence 2 or below, compared between arms by two-proportion z test, two-sided, alpha .05, against a stated 5% floor for calling the low end reachable at all. The direction is predicted and a null result is pre-committed as publishable. Why it matters beyond this instrument. The standard safety pattern for deploying a model into regulated work is triage: the system handles what it is confident about and escalates the rest to a person. That pattern requires the model to be able to report that it is unsure. If it cannot, the escalation step has nothing to act on. What this does not claim. The assessment items used here are not independently attested: zero items carry the 2-of-3 independent expert agreement the governing framework requires, which bounds every claim made from them. This is separate from the CRLN-TRG-001 competency crosswalk, whose ten canonical domains were SME-validated on 18 September 2026, when the third independent reviewer's ratings completed a unanimous 3 of 3. Self-reported confidence from a model is not the same object as a human's, and this study asks only whether the reported value moves with the elicitation.
Authors
- Joshua Webber (ORCID: https://orcid.org/0009-0005-2538-8333)
Institutions
- NIHR Research Delivery Network (GB)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22968474
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- preprint