CRLN-AIE-001 — Evaluating AI systems on clinical research competency against a human baseline: a reusable protocol and a first result
A reusable protocol for evaluating AI systems on clinical research competency, and a first result. Sponsors and CROs are deploying AI into GCP regulated work, and no published instrument exists for asking whether such a system is competent at those judgments. The component that is hardest to obtain is not the item bank but the matched human baseline, without which a model score is uninterpretable. First result. One frontier model (claude-sonnet-5) scored on 106 items that human learners had each attempted at least five times, giving 1,060 paired human decisions. Mean rubric score 3.632 for the model against 3.641 for humans, a difference of 0.008, with 96.2 percent item level agreement. Pass rates differed: 98.1 percent for the model against 91.3 percent for humans (z = 2.46, p = 0.014). The model is not better at the task so much as less variable. The primary finding is about the instrument, not the model. An item bank that cannot separate a competent human from a general purpose language model on 96 percent of its items cannot, without modification, certify either. We report this against our own instrument because the alternative is to publish a credential we cannot defend. The protocol is deliberately implementable without our platform. We invite other groups to run it on their own instruments and on other models, and to report results that contradict ours. Includes the evaluation harness, the edge function, and the human baseline snapshot.
Authors
- Joshua Webber (ORCID: https://orcid.org/0009-0005-2538-8333)
Institutions
- NIHR Research Delivery Network (GB)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-11
- DOI
- https://doi.org/10.5281/zenodo.22708752
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- preprint