Comparative performance of ChatGPT-5 mini and ChatGPT-5.3 in rhinologic emergency scenarios: an exploratory pilot study

Abstract Background With the growing use of Artificial Intelligence (AI)-based tools for health information, Large Language Models (LLMs) are increasingly used in clinical contexts. Although they show promising diagnostic performance in Ear, Nose, and Throat (ENT) emergencies, their reliability in rhinological scenarios remains uncertain due to lower treatment accuracy and the risk of hallucinations. Bu keşif niteliğindeki pilot çalışma, özellikle tanı doğruluğu, yönetim önerileri, aciliyet değerlendirmesi ve potansiyel olarak tehlikeli hatalara dikkat ederek, simüle edilmiş rinolojik acil durum senaryolarında ChatGPT-5 mini ve ChatGPT-5.3’ün performansını karşılaştırmayı amaçlamıştır. Methods Ten standardized rhinological emergency scenarios were independently submitted to ChatGPT-5 mini and ChatGPT-5.3 using the same standardized prompt. Model responses were independently evaluated by three otolaryngologists with more than five years of specialist experience. Diagnostic accuracy, treatment appropriateness, urgency assessment, and overall performance were rated using a 5-point Likert scale. Dangerous errors were assessed independently as present or absent and were classified at the scenario level when identified by at least two of the three evaluators. Paired model comparisons were performed using exact Wilcoxon signed-rank and McNemar tests, with Holm adjustment for multiple comparisons. Results Higher treatment and overall performance scores were observed for ChatGPT-5.3 than for GPT-5 mini. For treatment scores, the mean paired difference was 0.40 (BCa 95% CI: 0.33–0.47; exact p = 0.002; Holm-adjusted p = 0.006), while for overall performance the mean paired difference was 0.37 (BCa 95% CI: 0.17–0.57; exact p = 0.018; Holm-adjusted p = 0.036). Dangerous errors were identified in 4/10 (40%) scenarios for ChatGPT-5 mini and 2/10 (20%) for ChatGPT-5.3; this difference was not statistically significant (exact McNemar p = 0.500; Holm-adjusted p = 0.500). Diagnostic accuracy and urgency assessment showed no variability, with all responses receiving the maximum score, indicating a ceiling effect. Conclusion In this exploratory evaluation, differences between the two models were observed in treatment-related and overall performance, while no statistically significant difference was detected in dangerous-error rates. The ceiling effects in diagnostic accuracy and urgency assessment and the limited number of scenarios warrant cautious interpretation. These preliminary findings support the need for larger and more diverse evaluations before conclusions regarding the comparative clinical performance or safety of LLMs in rhinological emergencies can be drawn.

Authors

Institutions

Publication Details

Journal
The Egyptian Journal of Otolaryngology
Published
2026-09-12
DOI
https://doi.org/10.1186/s43163-026-01226-w
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Comparative performance of ChatGPT-5 mini and ChatGPT-5.3 in rhinologic emergency scenarios: an exploratory pilot study

Pınar Tekin
The Egyptian Journal of Otolaryngology
Artificial Intelligence in Healthcare and Education
article

Comparative performance of ChatGPT-5 mini and ChatGPT-5.3 in rhinologic emergency scenarios: an exploratory pilot study

Pınar Tekin
article en

Abstract

Abstract Background With the growing use of Artificial Intelligence (AI)-based tools for health information, Large Language Models (LLMs) are increasingly used in clinical contexts. Although they show promising diagnostic performance in Ear, Nose, and Throat (ENT) emergencies, their reliability in rhinological scenarios remains uncertain due to lower treatment accuracy and the risk of hallucinations. Bu keşif niteliğindeki pilot çalışma, özellikle tanı doğruluğu, yönetim önerileri, aciliyet değerlendirmesi ve potansiyel olarak tehlikeli hatalara dikkat ederek, simüle edilmiş rinolojik acil durum senaryolarında ChatGPT-5 mini ve ChatGPT-5.3’ün performansını karşılaştırmayı amaçlamıştır. Methods Ten standardized rhinological emergency scenarios were independently submitted to ChatGPT-5 mini and ChatGPT-5.3 using the same standardized prompt. Model responses were independently evaluated by three otolaryngologists with more than five years of specialist experience. Diagnostic accuracy, treatment appropriateness, urgency assessment, and overall performance were rated using a 5-point Likert scale. Dangerous errors were assessed independently as present or absent and were classified at the scenario level when identified by at least two of the three evaluators. Paired model comparisons were performed using exact Wilcoxon signed-rank and McNemar tests, with Holm adjustment for multiple comparisons. Results Higher treatment and overall performance scores were observed for ChatGPT-5.3 than for GPT-5 mini. For treatment scores, the mean paired difference was 0.40 (BCa 95% CI: 0.33–0.47; exact p = 0.002; Holm-adjusted p = 0.006), while for overall performance the mean paired difference was 0.37 (BCa 95% CI: 0.17–0.57; exact p = 0.018; Holm-adjusted p = 0.036). Dangerous errors were identified in 4/10 (40%) scenarios for ChatGPT-5 mini and 2/10 (20%) for ChatGPT-5.3; this difference was not statistically significant (exact McNemar p = 0.500; Holm-adjusted p = 0.500). Diagnostic accuracy and urgency assessment showed no variability, with all responses receiving the maximum score, indicating a ceiling effect. Conclusion In this exploratory evaluation, differences between the two models were observed in treatment-related and overall performance, while no statistically significant difference was detected in dangerous-error rates. The ceiling effects in diagnostic accuracy and urgency assessment and the limited number of scenarios warrant cautious interpretation. These preliminary findings support the need for larger and more diverse evaluations before conclusions regarding the comparative clinical performance or safety of LLMs in rhinological emergencies can be drawn.

The Egyptian Journal of OtolaryngologyVol. 42(1)
Malatya Devlet Hastanesi (TR)
Quality Education
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.