Evidence on Demand: OpenEvidence versus Established LLMs on the Orthopaedic In-Training Examination
Background Large language models (LLMs) are increasingly used in medical education, yet head-to-head comparisons of contemporary multimodal and clinically oriented (retrieval-grounded) models on orthopedic in-service examinations remain limited. Methods A cross-sectional comparative study was conducted using 265 multiple-choice questions from the publicly available 2014 Orthopaedic In-Training Examination (OITE). Questions were categorized by AAOS subspecialty and by cognitive taxonomy (T1 recall, T2 interpretation, T3 reasoning) and stratified into text-only (n=111) versus image-based items (n=164). ChatGPT-5 Plus and Gemini 2.5 Pro were evaluated on the full dataset, while OpenEvidence was evaluated on the text-only subset. Accuracy was scored against the official answer key and comparisons were done using McNemar and Cochran Q tests. Results Overall accuracy on the full 2014 OITE was 80.0% for Gemini 2.5 Pro and 78.1% for ChatGPT-5, with no statistically significant difference (p=0.59). Across the text-only subset, Gemini 2.5 Pro and OpenEvidence each scored 84.7% versus 81.1% for ChatGPT-5 (p=0.71), and OpenEvidence demonstrated strong T3 reasoning performance (82.7%). Subspecialty performance varied without statistically significant differences. LLMs demonstrated highest accuracy in basic science and the largest numerical discrepancy was in Oncology (Gemini 77.3% vs ChatGPT-5 54.5%). Conclusion A retrieval-grounded, evidence based LLM (OpenEvidence) performed comparably to multimodal LLMs on text-based orthopaedic in-training examination questions at or above senior resident PGY-5 benchmark-level performance with no significant differences in overall accuracy. These findings suggest that OpenEvidence may provide educational utility comparable to general-purpose models while offering source transparency.
Authors
- Zaamin B. Hussain (ORCID: https://orcid.org/0000-0003-3417-0123)
- Grant E. Garrigues (ORCID: https://orcid.org/0000-0002-8662-3131)
- Gregory P. Nicholson
- Lord J. Hyeamang (ORCID: https://orcid.org/0009-0007-5205-048X)
- Asim A. Khan (ORCID: https://orcid.org/0009-0007-8974-2332)
- Shaan S. Lalvani
- Ryan M. Lew
- Sam Pourarbab
Institutions
- Rush University Medical Center (US)
Publication Details
- Journal
- Journal of Orthopaedic Experience & Innovation
- Published
- 2026-09-26
- DOI
- https://doi.org/10.60118/001c.162882
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00