Evidence on Demand: OpenEvidence versus Established LLMs on the Orthopaedic In-Training Examination

Background Large language models (LLMs) are increasingly used in medical education, yet head-to-head comparisons of contemporary multimodal and clinically oriented (retrieval-grounded) models on orthopedic in-service examinations remain limited. Methods A cross-sectional comparative study was conducted using 265 multiple-choice questions from the publicly available 2014 Orthopaedic In-Training Examination (OITE). Questions were categorized by AAOS subspecialty and by cognitive taxonomy (T1 recall, T2 interpretation, T3 reasoning) and stratified into text-only (n=111) versus image-based items (n=164). ChatGPT-5 Plus and Gemini 2.5 Pro were evaluated on the full dataset, while OpenEvidence was evaluated on the text-only subset. Accuracy was scored against the official answer key and comparisons were done using McNemar and Cochran Q tests. Results Overall accuracy on the full 2014 OITE was 80.0% for Gemini 2.5 Pro and 78.1% for ChatGPT-5, with no statistically significant difference (p=0.59). Across the text-only subset, Gemini 2.5 Pro and OpenEvidence each scored 84.7% versus 81.1% for ChatGPT-5 (p=0.71), and OpenEvidence demonstrated strong T3 reasoning performance (82.7%). Subspecialty performance varied without statistically significant differences. LLMs demonstrated highest accuracy in basic science and the largest numerical discrepancy was in Oncology (Gemini 77.3% vs ChatGPT-5 54.5%). Conclusion A retrieval-grounded, evidence based LLM (OpenEvidence) performed comparably to multimodal LLMs on text-based orthopaedic in-training examination questions at or above senior resident PGY-5 benchmark-level performance with no significant differences in overall accuracy. These findings suggest that OpenEvidence may provide educational utility comparable to general-purpose models while offering source transparency.

Authors

Institutions

Publication Details

Journal
Journal of Orthopaedic Experience & Innovation
Published
2026-09-26
DOI
https://doi.org/10.60118/001c.162882
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evidence on Demand: OpenEvidence versus Established LLMs on the Orthopaedic In-Training Examination

Zaamin B. Hussain, Grant E. Garrigues, Gregory P. Nicholson, Lord J. Hyeamang et al.
Journal of Orthopaedic Experience & Innovation
Artificial Intelligence in Healthcare and Education
article

Evidence on Demand: OpenEvidence versus Established LLMs on the Orthopaedic In-Training Examination

Zaamin B. Hussain, Grant E. Garrigues, Gregory P. Nicholson, Lord J. Hyeamang, Asim A. Khan, Shaan S. Lalvani, Ryan M. Lew, Sam Pourarbab
article en

Abstract

Background Large language models (LLMs) are increasingly used in medical education, yet head-to-head comparisons of contemporary multimodal and clinically oriented (retrieval-grounded) models on orthopedic in-service examinations remain limited. Methods A cross-sectional comparative study was conducted using 265 multiple-choice questions from the publicly available 2014 Orthopaedic In-Training Examination (OITE). Questions were categorized by AAOS subspecialty and by cognitive taxonomy (T1 recall, T2 interpretation, T3 reasoning) and stratified into text-only (n=111) versus image-based items (n=164). ChatGPT-5 Plus and Gemini 2.5 Pro were evaluated on the full dataset, while OpenEvidence was evaluated on the text-only subset. Accuracy was scored against the official answer key and comparisons were done using McNemar and Cochran Q tests. Results Overall accuracy on the full 2014 OITE was 80.0% for Gemini 2.5 Pro and 78.1% for ChatGPT-5, with no statistically significant difference (p=0.59). Across the text-only subset, Gemini 2.5 Pro and OpenEvidence each scored 84.7% versus 81.1% for ChatGPT-5 (p=0.71), and OpenEvidence demonstrated strong T3 reasoning performance (82.7%). Subspecialty performance varied without statistically significant differences. LLMs demonstrated highest accuracy in basic science and the largest numerical discrepancy was in Oncology (Gemini 77.3% vs ChatGPT-5 54.5%). Conclusion A retrieval-grounded, evidence based LLM (OpenEvidence) performed comparably to multimodal LLMs on text-based orthopaedic in-training examination questions at or above senior resident PGY-5 benchmark-level performance with no significant differences in overall accuracy. These findings suggest that OpenEvidence may provide educational utility comparable to general-purpose models while offering source transparency.

Journal of Orthopaedic Experience & InnovationVol. 7(2)
Rush University Medical Center (US)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.