CRLN-PRE-001 — Pre-registration of two hypotheses on human-machine judgment

CRLN is a clinical research organisation, pre-registration is the norm we teach, and it is not one we had applied to our own work. This closes that gap before results arrive rather than after. It matters more than usual here: we have already published one correction, CRLN-HB-001 v0.2 having reported a human baseline that proved to be roughly 70% internal test accounts, with v0.3 reversing its headline conclusion. A group that has corrected itself once has a stronger obligation to fix its analysis plan in advance, not a weaker one. H1: human-model divergence is predictable from competency. CRLN-HMB-001 established that divergence VARIES, from humans +0.923 to models +0.377 across role-competency cells. It did not establish predictability. Pre-specified test: split-half correlation of per-cell differences, requiring 10 paired items per cell. At registration the best-covered cell holds 5. The grouping rule is fixed in advance, by role AND positional competency code jointly and never the bare code, because our own first analysis made that error and it reversed the finding. H2: a confident but wrong model recommendation degrades competent human judgement. Every deployment of AI into regulated work assumes a human-in-the-loop is a safeguard. If a confident wrong machine reliably talks competent people out of correct judgements, the safeguard is a rubber stamp on exactly the decisions where it matters most. Within-subject randomised design, three arms (no recommendation, sound recommendation, unsound recommendation), primary outcome pass rate in the unsound arm against control, threshold 200 decisions. At registration 144 items carry a gated recommendation and 3 are unsound, so the binding constraint is unsound items rather than learners. Learner protection is fixed in advance: the suggestion is labelled AI-generated and possibly wrong, rubric feedback still corrects, it never appears in a learner's first five decisions, and scoring is untouched. If any of those stops holding the experiment is switched off rather than adjusted. Commitments: neither hypothesis reported before its threshold, thresholds never lowered after seeing data, null results published with equal prominence, deviations reported as deviations. Licensed CC BY 4.0.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-15
DOI
https://doi.org/10.5281/zenodo.22760325
Citations
2
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

CRLN-PRE-001 — Pre-registration of two hypotheses on human-machine judgment

Joshua Webber
2 citations
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
preprint

CRLN-PRE-001 — Pre-registration of two hypotheses on human-machine judgment

Joshua Webber
preprint en
2 citations

Abstract

CRLN is a clinical research organisation, pre-registration is the norm we teach, and it is not one we had applied to our own work. This closes that gap before results arrive rather than after. It matters more than usual here: we have already published one correction, CRLN-HB-001 v0.2 having reported a human baseline that proved to be roughly 70% internal test accounts, with v0.3 reversing its headline conclusion. A group that has corrected itself once has a stronger obligation to fix its analysis plan in advance, not a weaker one. H1: human-model divergence is predictable from competency. CRLN-HMB-001 established that divergence VARIES, from humans +0.923 to models +0.377 across role-competency cells. It did not establish predictability. Pre-specified test: split-half correlation of per-cell differences, requiring 10 paired items per cell. At registration the best-covered cell holds 5. The grouping rule is fixed in advance, by role AND positional competency code jointly and never the bare code, because our own first analysis made that error and it reversed the finding. H2: a confident but wrong model recommendation degrades competent human judgement. Every deployment of AI into regulated work assumes a human-in-the-loop is a safeguard. If a confident wrong machine reliably talks competent people out of correct judgements, the safeguard is a rubber stamp on exactly the decisions where it matters most. Within-subject randomised design, three arms (no recommendation, sound recommendation, unsound recommendation), primary outcome pass rate in the unsound arm against control, threshold 200 decisions. At registration 144 items carry a gated recommendation and 3 are unsound, so the binding constraint is unsound items rather than learners. Learner protection is fixed in advance: the suggestion is labelled AI-generated and possibly wrong, rubric feedback still corrects, it never appears in a learner's first five decisions, and scoring is untouched. If any of those stops holding the experiment is switched off rather than adjusted. Commitments: neither hypothesis reported before its threshold, thresholds never lowered after seeing data, null results published with equal prominence, deviations reported as deviations. Licensed CC BY 4.0.

Zenodo (CERN European Organization for Nuclear Research)
NIHR Research Delivery Network (GB)
Peace, Justice and strong institutions
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.