Diagnostic Performance of a Generative Artificial Intelligence Model in Pelvic Ring and Acetabular Fractures: A Reliability and Validity Study
Aim: The objective of this study was to evaluate the effectiveness, accuracy, and consistency of an AI-powered chatbot (ChatGPT 4.0) in supporting clinical and radiological suspicions in the preliminary diagnosis of pelvic ring and acetabular fractures in situations where expert consultation or tertiary healthcare resources are unavailable. Material and Methods: A total of 100 antero-posterior (AP) pelvis X-ray images, 88 exhibiting pelvic or peri-acetabular fractures or instability, and all free from copyright restrictions, were presented to ChatGPT 4.0. In a scenario posing as a medical practitioner without any specialization, currently working in an emergency department in a rural hospital, and without any trauma or radiology specialist capable of evaluating the radiographic images, ChatGPT was asked to diagnose all x-rays, and then asked to confirm the diagnosis. Results: In only 43 (48.9%) of the 88 direct radiographs with fracture and instability findings, ChatGPT 4.0 was able to make a correct diagnosis, with a weak-to-minimal level of agreement between initial and confirmation answers. Furthermore, in 15 radiographs (17%), the fracture or instability was misdiagnosed as "solid," and in 30 radiographs (33.4%), while the presence of fracture or instability was detected, the definitive diagnosis was inaccurate. The pelvic ring and acetabular fractures that ChatGPT 4.0 correctly diagnosed at the highest rate were multiple fractures of the pelvic ring and acetabulum (80%), fractured dislocations of the acetabulum (72.7%), and isolated pubic ramus fractures (62.5%). ChatGPT 4.0 demonstrated the most inaccurate diagnostic performance for the following pelvic ring and acetabulum fractures, with a rate of 0%: femoral neck fractures, isolated iliac crest fractures, and vertically unstable pelvic ring injuries. Conclusion: In circumstances where there is a shortage of radiologists and orthopaedic specialists, such as in contexts where resources are limited or in rural emergency situations, reliance on AI-assisted chatbots such as ChatGPT 4.0 for fracture screening may result in diagnostic oversights or mismanagement. While the AI has been demonstrated to function as a supplementary tool to facilitate preliminary assessment, particularly in light of its optimal negative predictive value in non-fracture cases, its moderate false-negative rate (17%) necessitates a cautious approach.
Authors
- Batuhan Gencer (ORCID: https://orcid.org/0000-0003-0041-7378)
- İhsan Özdamar (ORCID: https://orcid.org/0000-0002-0685-9284)
- Ufuk Arzu (ORCID: https://orcid.org/0000-0002-6371-2211)
- Serdar Satılmış Orhan (ORCID: https://orcid.org/0000-0002-1358-5613)
- Turgut Dinçal
Institutions
- Marmara University (TR)
Publication Details
- Journal
- Anatolian Journal of Emergency Medicine
- Published
- 2026-09-24
- DOI
- https://doi.org/10.54996/anatolianjem.1772154
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00