An examiner-conditioned AI second marker for VR OSCEs
Abstract Objective Structured Clinical Examinations (OSCEs) are typically scored by a single examiner per station, exposing results to hard-to-audit examiner variability. We developed the Rater-Aware Verification Network (RAVEN), a multimodal Artificial Intelligence (AI) system fusing egocentric video, examiner verbalisations and marks, and virtual reality (VR) action logs as an examiner-conditioned second marker for paediatric OSCEs (retrospective evaluation; 120 students, 442 ratings, eight domains). A confidence-gated hybrid improved agreement with leave-one-examiner-out consensus on the pass/fail decision (AC1: +6.2 percentage points; p = 0.0004) and domain scores (mean AC2: +3.0 points, five of eight significant), with the largest gains on borderline-fail cases (+16.3 points of concordance with consensus). Comparing AI-inferred, rubric-specified, and examiner-articulated criteria revealed implicit practices absent from marking guidelines. Agreement is measured against a panel-derived reference rather than an external ground truth; the system is intended as an audit and flagging aid for human adjudication rather than an autonomous decision-maker.
Authors
- Sally Shiels
- Sajan B Patel (ORCID: https://orcid.org/0000-0002-8029-498X)
- Helen Higham (ORCID: https://orcid.org/0000-0001-5796-0377)
- Ashley Tomlinson
- Julia Alison Noble (ORCID: https://orcid.org/0000-0002-3060-3772)
- Nathan Gauge
- Harry Rogers (ORCID: https://orcid.org/0000-0003-3227-5677)
- James Thomas
- James Aylward
- Angela Feixue Wang
Institutions
- John Radcliffe Hospital (GB)
- University of Oxford (GB)
Publication Details
- Journal
- npj Digital Medicine
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1038/s41746-026-03336-4
- Primary Topic
- Innovations in Medical Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00