Impact of data origin on large language models’ radiologic interpretation: a proof-of-concept study on post-metabolic and bariatric surgery complications

Multimodal large language models (LLMs) show promise in medical imaging by integrating text and image data to help diagnose, improve consistency, and streamline workflows. They can approach specialist-level performance, support clinical decision-making, and assist in metabolic bariatric surgery (MBS) imaging, though applications remain early. Although LLMs perform well on exams, their real-world readiness still demands thorough validation before they can be safely adopted at scale. This study compared ChatGPT 5.2, Grok-4.1, and Gemini-3 Pro in diagnosing postoperative complications after MBS using 12 radiologic scenarios (six unpublished institutional cases and six PubMed-sourced cases). Each LLM received standardized prompts and case imaging. Model responses were compared with the reference diagnoses using top-1 concordance and top-3 inclusion as exploratory performance measures. Across 12 scenarios, ChatGPT 5.2 achieved the highest overall score (18/24), followed by Gemini-3 Pro (17/24) and Grok-4.1 (17/24), with no model achieving full concordance. All models demonstrated better performance in identifying structural and anatomical complications than functional or vascular complications. Published cases showed better observed diagnostic performance than previously unpublished cases; however, this difference should be interpreted cautiously because publicly available cases may have been represented in model training data, and therefore the observed advantage may reflect both recognition of well-documented clinical patterns and potential prior exposure to published case content. In this proof-of-concept evaluation, multimodal LLMs demonstrated potential to support differential diagnosis of postoperative MBS complications but showed variable performance, particularly for less common vascular and bowel-related complications. Higher observed scores for published than unpublished cases should be interpreted cautiously because prior exposure to publicly available case content during model training cannot be excluded and therefore cannot be considered evidence of superior diagnostic accuracy or generalizability. These findings support further evaluation using larger datasets of demonstrably unseen cases before clinical implementation.

Authors

Institutions

Publication Details

Journal
BMC Surgery
Published
2026-10-09
DOI
https://doi.org/10.1186/s12893-026-04265-5
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Impact of data origin on large language models’ radiologic interpretation: a proof-of-concept study on post-metabolic and bariatric surgery complications

Sonja Chiappetta, Mahsa Taherzadeh, Dániel Gerö, Shahab ShahabiShahmiri et al.
BMC Surgery
Artificial Intelligence in Healthcare and Education
article

Impact of data origin on large language models’ radiologic interpretation: a proof-of-concept study on post-metabolic and bariatric surgery complications

Sonja Chiappetta, Mahsa Taherzadeh, Dániel Gerö, Shahab ShahabiShahmiri, Panagiotis Laïnas, Ali Mousavimaleki, Mohammad Kermansaravi, Salvatore Tolone, Amir Hossein Davarpanah Jazi
article en

Abstract

Multimodal large language models (LLMs) show promise in medical imaging by integrating text and image data to help diagnose, improve consistency, and streamline workflows. They can approach specialist-level performance, support clinical decision-making, and assist in metabolic bariatric surgery (MBS) imaging, though applications remain early. Although LLMs perform well on exams, their real-world readiness still demands thorough validation before they can be safely adopted at scale. This study compared ChatGPT 5.2, Grok-4.1, and Gemini-3 Pro in diagnosing postoperative complications after MBS using 12 radiologic scenarios (six unpublished institutional cases and six PubMed-sourced cases). Each LLM received standardized prompts and case imaging. Model responses were compared with the reference diagnoses using top-1 concordance and top-3 inclusion as exploratory performance measures. Across 12 scenarios, ChatGPT 5.2 achieved the highest overall score (18/24), followed by Gemini-3 Pro (17/24) and Grok-4.1 (17/24), with no model achieving full concordance. All models demonstrated better performance in identifying structural and anatomical complications than functional or vascular complications. Published cases showed better observed diagnostic performance than previously unpublished cases; however, this difference should be interpreted cautiously because publicly available cases may have been represented in model training data, and therefore the observed advantage may reflect both recognition of well-documented clinical patterns and potential prior exposure to published case content. In this proof-of-concept evaluation, multimodal LLMs demonstrated potential to support differential diagnosis of postoperative MBS complications but showed variable performance, particularly for less common vascular and bowel-related complications. Higher observed scores for published than unpublished cases should be interpreted cautiously because prior exposure to publicly available case content during model training cannot be excluded and therefore cannot be considered evidence of superior diagnostic accuracy or generalizability. These findings support further evaluation using larger datasets of demonstrably unseen cases before clinical implementation.

BMC Surgery
European University Cyprus (CY), University of Zurich (CH), University Hospital of Zurich (CH), University of Campania "Luigi Vanvitelli" (IT), Metropolitan Hospital (GR)
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.