Diagnostic Performance of a Generative Artificial Intelligence Model in Pelvic Ring and Acetabular Fractures: A Reliability and Validity Study

Aim: The objective of this study was to evaluate the effectiveness, accuracy, and consistency of an AI-powered chatbot (ChatGPT 4.0) in supporting clinical and radiological suspicions in the preliminary diagnosis of pelvic ring and acetabular fractures in situations where expert consultation or tertiary healthcare resources are unavailable. Material and Methods: A total of 100 antero-posterior (AP) pelvis X-ray images, 88 exhibiting pelvic or peri-acetabular fractures or instability, and all free from copyright restrictions, were presented to ChatGPT 4.0. In a scenario posing as a medical practitioner without any specialization, currently working in an emergency department in a rural hospital, and without any trauma or radiology specialist capable of evaluating the radiographic images, ChatGPT was asked to diagnose all x-rays, and then asked to confirm the diagnosis. Results: In only 43 (48.9%) of the 88 direct radiographs with fracture and instability findings, ChatGPT 4.0 was able to make a correct diagnosis, with a weak-to-minimal level of agreement between initial and confirmation answers. Furthermore, in 15 radiographs (17%), the fracture or instability was misdiagnosed as "solid," and in 30 radiographs (33.4%), while the presence of fracture or instability was detected, the definitive diagnosis was inaccurate. The pelvic ring and acetabular fractures that ChatGPT 4.0 correctly diagnosed at the highest rate were multiple fractures of the pelvic ring and acetabulum (80%), fractured dislocations of the acetabulum (72.7%), and isolated pubic ramus fractures (62.5%). ChatGPT 4.0 demonstrated the most inaccurate diagnostic performance for the following pelvic ring and acetabulum fractures, with a rate of 0%: femoral neck fractures, isolated iliac crest fractures, and vertically unstable pelvic ring injuries. Conclusion: In circumstances where there is a shortage of radiologists and orthopaedic specialists, such as in contexts where resources are limited or in rural emergency situations, reliance on AI-assisted chatbots such as ChatGPT 4.0 for fracture screening may result in diagnostic oversights or mismanagement. While the AI has been demonstrated to function as a supplementary tool to facilitate preliminary assessment, particularly in light of its optimal negative predictive value in non-fracture cases, its moderate false-negative rate (17%) necessitates a cautious approach.

Authors

Institutions

Publication Details

Journal
Anatolian Journal of Emergency Medicine
Published
2026-09-24
DOI
https://doi.org/10.54996/anatolianjem.1772154
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Diagnostic Performance of a Generative Artificial Intelligence Model in Pelvic Ring and Acetabular Fractures: A Reliability and Validity Study

Batuhan Gencer, İhsan Özdamar, Ufuk Arzu, Serdar Satılmış Orhan et al.
Anatolian Journal of Emergency Medicine
Artificial Intelligence in Healthcare and Education
article

Diagnostic Performance of a Generative Artificial Intelligence Model in Pelvic Ring and Acetabular Fractures: A Reliability and Validity Study

Batuhan Gencer, İhsan Özdamar, Ufuk Arzu, Serdar Satılmış Orhan, Turgut Dinçal
article en

Abstract

Aim: The objective of this study was to evaluate the effectiveness, accuracy, and consistency of an AI-powered chatbot (ChatGPT 4.0) in supporting clinical and radiological suspicions in the preliminary diagnosis of pelvic ring and acetabular fractures in situations where expert consultation or tertiary healthcare resources are unavailable. Material and Methods: A total of 100 antero-posterior (AP) pelvis X-ray images, 88 exhibiting pelvic or peri-acetabular fractures or instability, and all free from copyright restrictions, were presented to ChatGPT 4.0. In a scenario posing as a medical practitioner without any specialization, currently working in an emergency department in a rural hospital, and without any trauma or radiology specialist capable of evaluating the radiographic images, ChatGPT was asked to diagnose all x-rays, and then asked to confirm the diagnosis. Results: In only 43 (48.9%) of the 88 direct radiographs with fracture and instability findings, ChatGPT 4.0 was able to make a correct diagnosis, with a weak-to-minimal level of agreement between initial and confirmation answers. Furthermore, in 15 radiographs (17%), the fracture or instability was misdiagnosed as "solid," and in 30 radiographs (33.4%), while the presence of fracture or instability was detected, the definitive diagnosis was inaccurate. The pelvic ring and acetabular fractures that ChatGPT 4.0 correctly diagnosed at the highest rate were multiple fractures of the pelvic ring and acetabulum (80%), fractured dislocations of the acetabulum (72.7%), and isolated pubic ramus fractures (62.5%). ChatGPT 4.0 demonstrated the most inaccurate diagnostic performance for the following pelvic ring and acetabulum fractures, with a rate of 0%: femoral neck fractures, isolated iliac crest fractures, and vertically unstable pelvic ring injuries. Conclusion: In circumstances where there is a shortage of radiologists and orthopaedic specialists, such as in contexts where resources are limited or in rural emergency situations, reliance on AI-assisted chatbots such as ChatGPT 4.0 for fracture screening may result in diagnostic oversights or mismanagement. While the AI has been demonstrated to function as a supplementary tool to facilitate preliminary assessment, particularly in light of its optimal negative predictive value in non-fracture cases, its moderate false-negative rate (17%) necessitates a cautious approach.

Anatolian Journal of Emergency MedicineVol. 9(3)
Marmara University (TR)
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.