Large Language Models in Multidisciplinary Decision-Making for Hepatopancreatobiliary Oncology: Retrospective Comparative Feasibility Study
Abstract Background Hepatopancreatobiliary (HPB) malignancies require complex treatment planning that often relies on multidisciplinary team (MDT) discussions. Large language models (LLMs) have recently been explored for clinical decision support, but their performance within real-world multidisciplinary decision environments remains unclear. In particular, the stability of LLM-generated recommendations—that is, whether a model produces the same answer when given the same clinical input—has rarely been examined. Objective This study aimed to evaluate the stability of treatment recommendations generated by contemporary LLMs when identical HPB cases are queried repeatedly, and their concordance with the treatment decisions reached at an institutional MDT conference. Methods This retrospective study included consecutive cases discussed at a single-center HPB MDT conference between September 1, 2024, and August 31, 2025. Standardized clinical case summaries derived from preconference documentation were provided to 4 LLMs (GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5) through their consumer web interfaces. Each model recommended a treatment among predefined MDT treatment options, and identical queries were repeated 4 times in separate sessions. Stability was quantified as the discordance rate relative to the initial response and, without privileging any single query, as the mean pairwise agreement and Fleiss κ across the 4 iterations. Concordance with MDT decisions was assessed using both the initial and modal responses, together with Cohen κ and class-wise F 1 -scores. Results A total of 107 MDT cases were analyzed. Stability differed significantly across models ( P =.01). Gemini 3 Pro showed the lowest discordance rate (mean 12.8%, SD 2.3%) and the highest reference-free agreement (Fleiss κ=0.737), whereas GPT-4o showed the highest discordance rate (mean 30.2%, SD 6.5%) and the lowest agreement (Fleiss κ=0.430). Concordance with MDT decisions ranged from 48.6% to 72.9% using the initial response and from 66.3% to 74.5% using the modal response, and the highest-performing model differed between the 2 definitions. Class-wise F 1 -score was consistently lower for surgery (0.400‐0.520) than for chemotherapy (0.621‐0.836). Complete discordance occurred in 17 of 107 (15.9%) cases and in none of the 31 anatomically unresectable cases (Fisher exact test, P =.003). Recurrent or on-treatment disease (adjusted odds ratio [OR] 5.40, 95% CI 1.65-17.68; P =.005), pancreatic tumor location (adjusted OR 7.37, 95% CI 2.03-26.78; P =.002), and low MDT agreement level (adjusted OR 10.33, 95% CI 1.54-69.38; P =.016) were independently associated with complete discordance. Conclusions LLM-generated treatment recommendations demonstrated moderate alignment with MDT decisions in HPB oncology. Importantly, response stability varied substantially across models, indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools. These findings suggest that LLMs may serve as a reasoning-support layer in MDT-like decision environments, but their response stability must be systematically characterized before clinical integration.
Authors
- Min Je Sung (ORCID: https://orcid.org/0000-0001-5395-8851)
- Incheon Kang (ORCID: https://orcid.org/0000-0003-4236-5094)
- Kwang Hyun Ko (ORCID: https://orcid.org/0000-0001-5168-1377)
- Ho Yeong Lim (ORCID: https://orcid.org/0000-0001-9325-2300)
- Chansik An (ORCID: https://orcid.org/0000-0002-0484-6658)
- Jung Ho Im (ORCID: https://orcid.org/0000-0002-3217-6444)
- Sung Jun Jo (ORCID: https://orcid.org/0000-0003-4638-652X)
- Hong Jae Chon (ORCID: https://orcid.org/0000-0002-6979-5812)
- Beodeul Kang (ORCID: https://orcid.org/0000-0001-5177-8937)
- Sung Hwan Lee (ORCID: https://orcid.org/0000-0003-3365-0096)
- Jeong‐Sik Yu (ORCID: https://orcid.org/0000-0002-8171-5838)
- Suk Pyo Shin (ORCID: https://orcid.org/0000-0002-5282-9174)
- Eui Hyuk Chong (ORCID: https://orcid.org/0000-0002-8998-7811)
- Jung Sun Kim (ORCID: https://orcid.org/0000-0003-1132-5217)
- Seok Jeong Yang (ORCID: https://orcid.org/0000-0001-6930-5978)
- Sujin Jang
- Chang-il Kwon
Publication Details
- Journal
- Journal of Medical Internet Research
- Published
- 2026-09-17
- DOI
- https://doi.org/10.2196/101137
- Primary Topic
- Cholangiocarcinoma and Gallbladder Cancer Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00