CRLN-AIE-001 v0.2 — What an instrument with a measured ceiling can and cannot tell you about frontier models

Version 0.2 of the CRLN AI competency evaluation protocol, correcting the central result of v0.1. Correction. v0.1 compared an unweighted model mean against a response-weighted human mean and concluded that a frontier model and competent practitioners were indistinguishable. Those are different quantities, differing by 0.093 on the same data because popular items are easier items. Compared like with like the conclusion reverses. Method. Seven models across three vendors (Anthropic, Google, OpenAI) evaluated on 114 clinical trial conduct decisions that human practitioners had each attempted at least five times. Five presentations per item with options re-randomized, a fixed-order control condition to separate position sensitivity from retest noise, and every one of 3,990 responses retained with the option permutation as presented and the vendor-reported model version. Results. The instrument discriminates. Two models are statistically indistinguishable from practitioners, four sit modestly above, one well above. Compliance depended on the order the options were presented in on 1 to 12 items of 114, and that count tracks capability. Position sensitivity accounts for roughly 80 percent of within-item variation in one model and 65 percent in another; to our knowledge this is the first partition of position and retest components within a single MCQA benchmark design. Reported against ourselves. The instrument cannot currently detect whether models flip between compliant and non-compliant, because its own ceiling sits above the range where that effect is detectable. Twenty-three items authored specifically to lower that ceiling did not work. The author's own responses have been excluded from the human baseline. The human comparison group is 52 practitioners, mostly early career. The full response dataset accompanies this version.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-12
DOI
https://doi.org/10.5281/zenodo.22728905
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

CRLN-AIE-001 v0.2 — What an instrument with a measured ceiling can and cannot tell you about frontier models

Joshua Webber
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
article

CRLN-AIE-001 v0.2 — What an instrument with a measured ceiling can and cannot tell you about frontier models

Joshua Webber
article en

Abstract

Version 0.2 of the CRLN AI competency evaluation protocol, correcting the central result of v0.1. Correction. v0.1 compared an unweighted model mean against a response-weighted human mean and concluded that a frontier model and competent practitioners were indistinguishable. Those are different quantities, differing by 0.093 on the same data because popular items are easier items. Compared like with like the conclusion reverses. Method. Seven models across three vendors (Anthropic, Google, OpenAI) evaluated on 114 clinical trial conduct decisions that human practitioners had each attempted at least five times. Five presentations per item with options re-randomized, a fixed-order control condition to separate position sensitivity from retest noise, and every one of 3,990 responses retained with the option permutation as presented and the vendor-reported model version. Results. The instrument discriminates. Two models are statistically indistinguishable from practitioners, four sit modestly above, one well above. Compliance depended on the order the options were presented in on 1 to 12 items of 114, and that count tracks capability. Position sensitivity accounts for roughly 80 percent of within-item variation in one model and 65 percent in another; to our knowledge this is the first partition of position and retest components within a single MCQA benchmark design. Reported against ourselves. The instrument cannot currently detect whether models flip between compliant and non-compliant, because its own ceiling sits above the range where that effect is detectable. Twenty-three items authored specifically to lower that ceiling did not work. The author's own responses have been excluded from the human baseline. The human comparison group is 52 practitioners, mostly early career. The full response dataset accompanies this version.

Zenodo (CERN European Organization for Nuclear Research)
NIHR Research Delivery Network (GB)
Reduced inequalities
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

CRLN-AIE-001 v0.2 — What an instrument with a measured ceiling can and cannot tell you about frontier models — Joshua Webber · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS