Large language models and specialty certification examinations in obstetrics and gynecology: a comparative performance analysis
Large language models (LLMs) are becoming integral to medical fields, yet comprehensive evaluations of their expertise in specific domains such as obstetrics and gynecology remain limited. The aim of this study was to assess and compare the observed performance and accuracy of a pragmatic sample of widely used LLMs—ChatGPT 5.1, Gemini Flash 2.5, Gemini Pro 3.0, Claude Sonnet 4.5, Microsoft Copilot, and Llama 4 Maverick—to determine their utility as educational tools and adjunctive knowledge-assistance systems. We submitted 355 single-best-answer questions from three recent sessions of the Polish Specialty Certificate Examination in Obstetrics and Gynecology to the models. Questions were translated into English, and model performance was evaluated based on overall accuracy and the difficulty index of the questions. The results demonstrated that while all models surpassed the 60% passing threshold required for human specialists, significant performance disparities were observed. Gemini Pro 3.0 achieved superior results, consistently exceeding 90% accuracy and maintaining stability regardless of question complexity. Conversely, descriptive session-stratified analyses indicated lower accuracy for Gemini Flash 2.5, Microsoft Copilot, and Llama 4 Maverick on difficult questions. These findings suggest that while LLMs hold promise as educational and knowledge-assistance tools, their performance variability necessitates rigorous prospective validation and expert oversight before any clinical deployment.
Authors
- Paweł Jan Stanirowski (ORCID: https://orcid.org/0000-0001-7445-7546)
- Franciszek Ługowski (ORCID: https://orcid.org/0000-0001-6952-4927)
- Anna Owczarek
Institutions
- Medical University of Warsaw (PL)
Publication Details
- Journal
- Archives of Gynecology and Obstetrics
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1007/s00404-026-08572-3
- Primary Topic
- Innovations in Medical Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00