Diagnostic performance of GPT-5.2 ınstant for pediatric elbow fracture detection on radiographs: a prospective single-center diagnostic accuracy study

Abstract Background Multimodal large language models can interpret medical images, but their performance for pediatric elbow radiographs remains uncertain. We evaluated the diagnostic performance of GPT-5.2 Instant as accessed through the ChatGPT web interface during the defined study period. Methods In this prospective, single-center diagnostic accuracy study, consecutive patients aged 0–17 years presenting with acute elbow trauma between December 2025 and May 2026 were included. De-identified anteroposterior and lateral radiographs exported from the institutional PACS as PDF files were evaluated by GPT-5.2 Instant through the standard ChatGPT web interface, once per case, using a fixed prompt and no task-specific fine-tuning. The final adjudicated composite reference standard was based on concordant assessments by two emergency physician readers; discordant cases underwent independent radiologist review and adjudication. Diagnostic performance was summarized with 95% confidence intervals. Results Among 412 patients, the reference-standard fracture prevalence was 10.0% (41/412). GPT-5.2 Instant classified 74 cases (18.0%) as fracture-positive. It correctly identified 23 of 41 fractures and missed 18. Sensitivity was 56.10% (95% CI, 39.75–71.53), specificity 86.25% (82.32–89.59), positive likelihood ratio 4.08 (2.81–5.92), negative likelihood ratio 0.51 (0.36–0.72), positive predictive value 31.08% (21.61–42.46), negative predictive value 94.67% (91.71–96.62), and overall accuracy 83.25% (79.29–86.73). Agreement with the reference standard was limited (κ = 0.312). Conclusions Under the reported access and input conditions, GPT-5.2 Instant demonstrated limited sensitivity and agreement for pediatric elbow fracture detection, insufficient to support independent diagnostic or rule-out use. Clinician–AI collaboration was not evaluated and requires prospective investigation.

Authors

Institutions

Publication Details

Journal
BMC Medical Imaging
Published
2026-09-09
DOI
https://doi.org/10.1186/s12880-026-02788-0
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Diagnostic performance of GPT-5.2 ınstant for pediatric elbow fracture detection on radiographs: a prospective single-center diagnostic accuracy study

Ramazan Sami Aktaş, Mahmut Şahin, Osman Taş, Mehmet Şirin Büyükkaya et al.
BMC Medical Imaging
Artificial Intelligence in Healthcare and Education
article

Diagnostic performance of GPT-5.2 ınstant for pediatric elbow fracture detection on radiographs: a prospective single-center diagnostic accuracy study

Ramazan Sami Aktaş, Mahmut Şahin, Osman Taş, Mehmet Şirin Büyükkaya, Mehmet Yorgun
article en

Abstract

Abstract Background Multimodal large language models can interpret medical images, but their performance for pediatric elbow radiographs remains uncertain. We evaluated the diagnostic performance of GPT-5.2 Instant as accessed through the ChatGPT web interface during the defined study period. Methods In this prospective, single-center diagnostic accuracy study, consecutive patients aged 0–17 years presenting with acute elbow trauma between December 2025 and May 2026 were included. De-identified anteroposterior and lateral radiographs exported from the institutional PACS as PDF files were evaluated by GPT-5.2 Instant through the standard ChatGPT web interface, once per case, using a fixed prompt and no task-specific fine-tuning. The final adjudicated composite reference standard was based on concordant assessments by two emergency physician readers; discordant cases underwent independent radiologist review and adjudication. Diagnostic performance was summarized with 95% confidence intervals. Results Among 412 patients, the reference-standard fracture prevalence was 10.0% (41/412). GPT-5.2 Instant classified 74 cases (18.0%) as fracture-positive. It correctly identified 23 of 41 fractures and missed 18. Sensitivity was 56.10% (95% CI, 39.75–71.53), specificity 86.25% (82.32–89.59), positive likelihood ratio 4.08 (2.81–5.92), negative likelihood ratio 0.51 (0.36–0.72), positive predictive value 31.08% (21.61–42.46), negative predictive value 94.67% (91.71–96.62), and overall accuracy 83.25% (79.29–86.73). Agreement with the reference standard was limited (κ = 0.312). Conclusions Under the reported access and input conditions, GPT-5.2 Instant demonstrated limited sensitivity and agreement for pediatric elbow fracture detection, insufficient to support independent diagnostic or rule-out use. Clinician–AI collaboration was not evaluated and requires prospective investigation.

BMC Medical Imaging
Education Training And Research (US), Sağlık Bilimleri Üniversitesi (TR), University of Health Sciences Antigua (AG)
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.