CRLN-AIE-001 — Evaluating AI systems on clinical research competency against a human baseline: a reusable protocol and a first result

A reusable protocol for evaluating AI systems on clinical research competency, and a first result. Sponsors and CROs are deploying AI into GCP regulated work, and no published instrument exists for asking whether such a system is competent at those judgments. The component that is hardest to obtain is not the item bank but the matched human baseline, without which a model score is uninterpretable. First result. One frontier model (claude-sonnet-5) scored on 106 items that human learners had each attempted at least five times, giving 1,060 paired human decisions. Mean rubric score 3.632 for the model against 3.641 for humans, a difference of 0.008, with 96.2 percent item level agreement. Pass rates differed: 98.1 percent for the model against 91.3 percent for humans (z = 2.46, p = 0.014). The model is not better at the task so much as less variable. The primary finding is about the instrument, not the model. An item bank that cannot separate a competent human from a general purpose language model on 96 percent of its items cannot, without modification, certify either. We report this against our own instrument because the alternative is to publish a credential we cannot defend. The protocol is deliberately implementable without our platform. We invite other groups to run it on their own instruments and on other models, and to report results that contradict ours. Includes the evaluation harness, the edge function, and the human baseline snapshot.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-11
DOI
https://doi.org/10.5281/zenodo.22708752
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

CRLN-AIE-001 — Evaluating AI systems on clinical research competency against a human baseline: a reusable protocol and a first result

Joshua Webber
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
preprint

CRLN-AIE-001 — Evaluating AI systems on clinical research competency against a human baseline: a reusable protocol and a first result

Joshua Webber
preprint en

Abstract

A reusable protocol for evaluating AI systems on clinical research competency, and a first result. Sponsors and CROs are deploying AI into GCP regulated work, and no published instrument exists for asking whether such a system is competent at those judgments. The component that is hardest to obtain is not the item bank but the matched human baseline, without which a model score is uninterpretable. First result. One frontier model (claude-sonnet-5) scored on 106 items that human learners had each attempted at least five times, giving 1,060 paired human decisions. Mean rubric score 3.632 for the model against 3.641 for humans, a difference of 0.008, with 96.2 percent item level agreement. Pass rates differed: 98.1 percent for the model against 91.3 percent for humans (z = 2.46, p = 0.014). The model is not better at the task so much as less variable. The primary finding is about the instrument, not the model. An item bank that cannot separate a competent human from a general purpose language model on 96 percent of its items cannot, without modification, certify either. We report this against our own instrument because the alternative is to publish a credential we cannot defend. The protocol is deliberately implementable without our platform. We invite other groups to run it on their own instruments and on other models, and to report results that contradict ours. Includes the evaluation harness, the edge function, and the human baseline snapshot.

Zenodo (CERN European Organization for Nuclear Research)
NIHR Research Delivery Network (GB)
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

CRLN-AIE-001 — Evaluating AI systems on clinical research competency against a human baseline: a reusable protocol and a first result — Joshua Webber · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS