Large language models and specialty certification examinations in obstetrics and gynecology: a comparative performance analysis

Large language models (LLMs) are becoming integral to medical fields, yet comprehensive evaluations of their expertise in specific domains such as obstetrics and gynecology remain limited. The aim of this study was to assess and compare the observed performance and accuracy of a pragmatic sample of widely used LLMs—ChatGPT 5.1, Gemini Flash 2.5, Gemini Pro 3.0, Claude Sonnet 4.5, Microsoft Copilot, and Llama 4 Maverick—to determine their utility as educational tools and adjunctive knowledge-assistance systems. We submitted 355 single-best-answer questions from three recent sessions of the Polish Specialty Certificate Examination in Obstetrics and Gynecology to the models. Questions were translated into English, and model performance was evaluated based on overall accuracy and the difficulty index of the questions. The results demonstrated that while all models surpassed the 60% passing threshold required for human specialists, significant performance disparities were observed. Gemini Pro 3.0 achieved superior results, consistently exceeding 90% accuracy and maintaining stability regardless of question complexity. Conversely, descriptive session-stratified analyses indicated lower accuracy for Gemini Flash 2.5, Microsoft Copilot, and Llama 4 Maverick on difficult questions. These findings suggest that while LLMs hold promise as educational and knowledge-assistance tools, their performance variability necessitates rigorous prospective validation and expert oversight before any clinical deployment.

Authors

Institutions

Publication Details

Journal
Archives of Gynecology and Obstetrics
Published
2026-09-16
DOI
https://doi.org/10.1007/s00404-026-08572-3
Primary Topic
Innovations in Medical Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Large language models and specialty certification examinations in obstetrics and gynecology: a comparative performance analysis

Paweł Jan Stanirowski, Franciszek Ługowski, Anna Owczarek
Archives of Gynecology and Obstetrics
Innovations in Medical Education
article

Large language models and specialty certification examinations in obstetrics and gynecology: a comparative performance analysis

Paweł Jan Stanirowski, Franciszek Ługowski, Anna Owczarek
article en

Abstract

Large language models (LLMs) are becoming integral to medical fields, yet comprehensive evaluations of their expertise in specific domains such as obstetrics and gynecology remain limited. The aim of this study was to assess and compare the observed performance and accuracy of a pragmatic sample of widely used LLMs—ChatGPT 5.1, Gemini Flash 2.5, Gemini Pro 3.0, Claude Sonnet 4.5, Microsoft Copilot, and Llama 4 Maverick—to determine their utility as educational tools and adjunctive knowledge-assistance systems. We submitted 355 single-best-answer questions from three recent sessions of the Polish Specialty Certificate Examination in Obstetrics and Gynecology to the models. Questions were translated into English, and model performance was evaluated based on overall accuracy and the difficulty index of the questions. The results demonstrated that while all models surpassed the 60% passing threshold required for human specialists, significant performance disparities were observed. Gemini Pro 3.0 achieved superior results, consistently exceeding 90% accuracy and maintaining stability regardless of question complexity. Conversely, descriptive session-stratified analyses indicated lower accuracy for Gemini Flash 2.5, Microsoft Copilot, and Llama 4 Maverick on difficult questions. These findings suggest that while LLMs hold promise as educational and knowledge-assistance tools, their performance variability necessitates rigorous prospective validation and expert oversight before any clinical deployment.

Archives of Gynecology and Obstetrics
Medical University of Warsaw (PL)
Quality Education
Openalex Percentile: Top 9%
Innovations in Medical Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Large language models and specialty certification examinations in obstetrics and gynecology: a comparative performance analysis — Paweł Jan Stanirowski, Franciszek Ługowski, et al. · Archives of Gynecology and Obstetrics (2026) | TGRS Research Map | TGRS