Concordance between GPT-4 and a multidisciplinary tumor board in pancreatic cancer: A prospective pilot study

Abstract Background Large language models (LLMs) such as GPT-4 are being evaluated for their use as supportive tools in oncological treatment planning. However, in pancreatic cancer, current studies are confined to predefined question–answer formats, while studies specifically investigating real-world scenarios that benchmark LLM performance against multidisciplinary tumor board (MDT) decisions are lacking. Methods This prospective comparative analysis evaluated treatment and diagnostic recommendations for patients with newly diagnosed or suspected pancreatic cancer between an MDT and GPT-4. Using MDT referrals, clinical data were entered into a clinical data matrix and submitted to GPT-4 for therapeutic and diagnostic recommendations. Outputs were assessed before and after additional prompting with 41 high-ranking abstracts relevant to pancreatic cancer care. The primary endpoint was the concordance of recommendations between the MDT and GPT-4 before and after literature-based prompting. Results Between September 1, 2024 and March 31, 2025, 45 patients were enrolled. The overall concordance rate between the MDT and GPT-4 was 73.3% (κ = 0.64, p < 0.0001) and did not improve following literature prompting. Discordance most often occurred in complex clinical scenarios. Concordance was highest in cases of metastatic disease (90.0%) and in neoadjuvant settings (90.0%) while it was lowest in patients requiring additional diagnostic workup (50.0%). Conclusions GPT-4 demonstrated substantial agreement with MDT recommendations in patients with newly diagnosed or suspected pancreatic cancer. However, specific abstract prompting did not enhance the rate of concordance and GPT-4’s limitations in individualized or complex contexts underscore the need for a cautious future integration into oncologic workflows.

Authors

Institutions

Publication Details

Journal
Langenbeck s Archives of Surgery
Published
2026-09-17
DOI
https://doi.org/10.1007/s00423-026-04242-9
Primary Topic
Pancreatic and Hepatic Oncology Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Concordance between GPT-4 and a multidisciplinary tumor board in pancreatic cancer: A prospective pilot study

Thilo Hackert, Antonie Willner, Mara Göetz, Anna Nießen et al.
Langenbeck s Archives of Surgery
Pancreatic and Hepatic Oncology Research
article

Concordance between GPT-4 and a multidisciplinary tumor board in pancreatic cancer: A prospective pilot study

Thilo Hackert, Antonie Willner, Mara Göetz, Anna Nießen, Jan Bardenhagen, Felix Nickel, Andreas Brandl, Thilo Welsch, Kürsat Kirkgöz, Fiete Gehrisch, Marianne Sinn, Faik G. Uzunoglu
article en

Abstract

Abstract Background Large language models (LLMs) such as GPT-4 are being evaluated for their use as supportive tools in oncological treatment planning. However, in pancreatic cancer, current studies are confined to predefined question–answer formats, while studies specifically investigating real-world scenarios that benchmark LLM performance against multidisciplinary tumor board (MDT) decisions are lacking. Methods This prospective comparative analysis evaluated treatment and diagnostic recommendations for patients with newly diagnosed or suspected pancreatic cancer between an MDT and GPT-4. Using MDT referrals, clinical data were entered into a clinical data matrix and submitted to GPT-4 for therapeutic and diagnostic recommendations. Outputs were assessed before and after additional prompting with 41 high-ranking abstracts relevant to pancreatic cancer care. The primary endpoint was the concordance of recommendations between the MDT and GPT-4 before and after literature-based prompting. Results Between September 1, 2024 and March 31, 2025, 45 patients were enrolled. The overall concordance rate between the MDT and GPT-4 was 73.3% (κ = 0.64, p < 0.0001) and did not improve following literature prompting. Discordance most often occurred in complex clinical scenarios. Concordance was highest in cases of metastatic disease (90.0%) and in neoadjuvant settings (90.0%) while it was lowest in patients requiring additional diagnostic workup (50.0%). Conclusions GPT-4 demonstrated substantial agreement with MDT recommendations in patients with newly diagnosed or suspected pancreatic cancer. However, specific abstract prompting did not enhance the rate of concordance and GPT-4’s limitations in individualized or complex contexts underscore the need for a cautious future integration into oncologic workflows.

Langenbeck s Archives of SurgeryVol. 411(1)
Universität Hamburg (DE), University Medical Center Hamburg-Eppendorf (DE), Krankenhaus Nordwest (DE)
Quality Education
Openalex Percentile: Top 14%
Pancreatic and Hepatic Oncology Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.