CRLN-AIE-001 v0.2 — What an instrument with a measured ceiling can and cannot tell you about frontier models
Version 0.2 of the CRLN AI competency evaluation protocol, correcting the central result of v0.1. Correction. v0.1 compared an unweighted model mean against a response-weighted human mean and concluded that a frontier model and competent practitioners were indistinguishable. Those are different quantities, differing by 0.093 on the same data because popular items are easier items. Compared like with like the conclusion reverses. Method. Seven models across three vendors (Anthropic, Google, OpenAI) evaluated on 114 clinical trial conduct decisions that human practitioners had each attempted at least five times. Five presentations per item with options re-randomized, a fixed-order control condition to separate position sensitivity from retest noise, and every one of 3,990 responses retained with the option permutation as presented and the vendor-reported model version. Results. The instrument discriminates. Two models are statistically indistinguishable from practitioners, four sit modestly above, one well above. Compliance depended on the order the options were presented in on 1 to 12 items of 114, and that count tracks capability. Position sensitivity accounts for roughly 80 percent of within-item variation in one model and 65 percent in another; to our knowledge this is the first partition of position and retest components within a single MCQA benchmark design. Reported against ourselves. The instrument cannot currently detect whether models flip between compliant and non-compliant, because its own ceiling sits above the range where that effect is detectable. Twenty-three items authored specifically to lower that ceiling did not work. The author's own responses have been excluded from the human baseline. The human comparison group is 52 practitioners, mostly early career. The full response dataset accompanies this version.
Authors
- Joshua Webber (ORCID: https://orcid.org/0009-0005-2538-8333)
Institutions
- NIHR Research Delivery Network (GB)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-12
- DOI
- https://doi.org/10.5281/zenodo.22728905
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00