Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions

Aim: In large language model (LLM) examination studies, models answer the same items repeatedly, so observations are clustered and conventional intervals overstate precision. We aimed to measure the accuracy of seven LLMs on a subspecialty-level anaesthesiology question bank and to determine which between-model and subgroup differences remain distinguishable under a cluster-respecting analysis.Materials and Methods: A 120-item, five-option multiple-choice bank in the style of the Anaesthesiology and Reanimation Subspecialty Examination, written by the authors and reviewed against current guidelines and textbooks, was administered to seven LLMs in five runs at temperature 0, with option order re-randomised per run (4,200 responses from 120 unique items). Accuracy is reported with cluster bootstrap 95% confidence intervals (CIs) resampling items (4,000 replicates) and naive Wilson intervals for comparison; subgroup analyses are exploratory.Results: Pooled accuracy was 80.1% (cluster-robust 95% CI 75.8-83.9; naive Wilson 78.8-81.3). Between-model differences were large and robust, spanning 60.8% (54.0-67.5) to 94.0% (90.3-97.0) with non-overlapping extremes. Cluster-robust intervals were wider in 25 of 26 estimates (median factor 2.1, range 0.8-4.0). Most subgroup comparisons were not distinguishable, intervals overlapping substantially: vignette (75.4%, 62.5-86.5) versus non-vignette (80.9%, 76.6-84.9) and guideline-dependent (73.0%, 64.6-80.7) versus other items (80.9%, 76.6-84.9). Only contrasts between weakest and strongest domains persisted. Between-run standard deviation was 1.26-2.80 points.Conclusion: Between-model differences and run-to-run instability are robust; commonly emphasised domain and item-characteristic differences are mostly not distinguishable once clustering is respected. Benchmarks of 120 items can separate models whose accuracies differ widely but cannot reliably localise their weaknesses.

Authors

Institutions

Publication Details

Journal
Journal of Contemporary Medicine
Published
2026-09-29
DOI
https://doi.org/10.16899/jcm.2008162
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions

Sait Ramazan GÜLBAY, Muhammed Nezih KOÇ
Journal of Contemporary Medicine
Artificial Intelligence in Healthcare and Education
article

Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions

Sait Ramazan GÜLBAY, Muhammed Nezih KOÇ
article en

Abstract

Aim: In large language model (LLM) examination studies, models answer the same items repeatedly, so observations are clustered and conventional intervals overstate precision. We aimed to measure the accuracy of seven LLMs on a subspecialty-level anaesthesiology question bank and to determine which between-model and subgroup differences remain distinguishable under a cluster-respecting analysis.Materials and Methods: A 120-item, five-option multiple-choice bank in the style of the Anaesthesiology and Reanimation Subspecialty Examination, written by the authors and reviewed against current guidelines and textbooks, was administered to seven LLMs in five runs at temperature 0, with option order re-randomised per run (4,200 responses from 120 unique items). Accuracy is reported with cluster bootstrap 95% confidence intervals (CIs) resampling items (4,000 replicates) and naive Wilson intervals for comparison; subgroup analyses are exploratory.Results: Pooled accuracy was 80.1% (cluster-robust 95% CI 75.8-83.9; naive Wilson 78.8-81.3). Between-model differences were large and robust, spanning 60.8% (54.0-67.5) to 94.0% (90.3-97.0) with non-overlapping extremes. Cluster-robust intervals were wider in 25 of 26 estimates (median factor 2.1, range 0.8-4.0). Most subgroup comparisons were not distinguishable, intervals overlapping substantially: vignette (75.4%, 62.5-86.5) versus non-vignette (80.9%, 76.6-84.9) and guideline-dependent (73.0%, 64.6-80.7) versus other items (80.9%, 76.6-84.9). Only contrasts between weakest and strongest domains persisted. Between-run standard deviation was 1.26-2.80 points.Conclusion: Between-model differences and run-to-run instability are robust; commonly emphasised domain and item-characteristic differences are mostly not distinguishable once clustering is respected. Benchmarks of 120 items can separate models whose accuracies differ widely but cannot reliably localise their weaknesses.

Journal of Contemporary MedicineVol. 16(5)
Necmettin Erbakan University (TR), Konya City Hospital (TR)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.