An examiner-conditioned AI second marker for VR OSCEs

Abstract Objective Structured Clinical Examinations (OSCEs) are typically scored by a single examiner per station, exposing results to hard-to-audit examiner variability. We developed the Rater-Aware Verification Network (RAVEN), a multimodal Artificial Intelligence (AI) system fusing egocentric video, examiner verbalisations and marks, and virtual reality (VR) action logs as an examiner-conditioned second marker for paediatric OSCEs (retrospective evaluation; 120 students, 442 ratings, eight domains). A confidence-gated hybrid improved agreement with leave-one-examiner-out consensus on the pass/fail decision (AC1: +6.2 percentage points; p = 0.0004) and domain scores (mean AC2: +3.0 points, five of eight significant), with the largest gains on borderline-fail cases (+16.3 points of concordance with consensus). Comparing AI-inferred, rubric-specified, and examiner-articulated criteria revealed implicit practices absent from marking guidelines. Agreement is measured against a panel-derived reference rather than an external ground truth; the system is intended as an audit and flagging aid for human adjudication rather than an autonomous decision-maker.

Authors

Institutions

Publication Details

Journal
npj Digital Medicine
Published
2026-10-07
DOI
https://doi.org/10.1038/s41746-026-03336-4
Primary Topic
Innovations in Medical Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

An examiner-conditioned AI second marker for VR OSCEs

Sally Shiels, Sajan B Patel, Helen Higham, Ashley Tomlinson et al.
npj Digital Medicine
Innovations in Medical Education
article

An examiner-conditioned AI second marker for VR OSCEs

Sally Shiels, Sajan B Patel, Helen Higham, Ashley Tomlinson, Julia Alison Noble, Nathan Gauge, Harry Rogers, James Thomas, James Aylward, Angela Feixue Wang
article en

Abstract

Abstract Objective Structured Clinical Examinations (OSCEs) are typically scored by a single examiner per station, exposing results to hard-to-audit examiner variability. We developed the Rater-Aware Verification Network (RAVEN), a multimodal Artificial Intelligence (AI) system fusing egocentric video, examiner verbalisations and marks, and virtual reality (VR) action logs as an examiner-conditioned second marker for paediatric OSCEs (retrospective evaluation; 120 students, 442 ratings, eight domains). A confidence-gated hybrid improved agreement with leave-one-examiner-out consensus on the pass/fail decision (AC1: +6.2 percentage points; p = 0.0004) and domain scores (mean AC2: +3.0 points, five of eight significant), with the largest gains on borderline-fail cases (+16.3 points of concordance with consensus). Comparing AI-inferred, rubric-specified, and examiner-articulated criteria revealed implicit practices absent from marking guidelines. Agreement is measured against a panel-derived reference rather than an external ground truth; the system is intended as an audit and flagging aid for human adjudication rather than an autonomous decision-maker.

npj Digital Medicine
John Radcliffe Hospital (GB), University of Oxford (GB)
Openalex Percentile: Top 9%
Innovations in Medical Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.