Large Language Models in Multidisciplinary Decision-Making for Hepatopancreatobiliary Oncology: Retrospective Comparative Feasibility Study

Abstract Background Hepatopancreatobiliary (HPB) malignancies require complex treatment planning that often relies on multidisciplinary team (MDT) discussions. Large language models (LLMs) have recently been explored for clinical decision support, but their performance within real-world multidisciplinary decision environments remains unclear. In particular, the stability of LLM-generated recommendations—that is, whether a model produces the same answer when given the same clinical input—has rarely been examined. Objective This study aimed to evaluate the stability of treatment recommendations generated by contemporary LLMs when identical HPB cases are queried repeatedly, and their concordance with the treatment decisions reached at an institutional MDT conference. Methods This retrospective study included consecutive cases discussed at a single-center HPB MDT conference between September 1, 2024, and August 31, 2025. Standardized clinical case summaries derived from preconference documentation were provided to 4 LLMs (GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5) through their consumer web interfaces. Each model recommended a treatment among predefined MDT treatment options, and identical queries were repeated 4 times in separate sessions. Stability was quantified as the discordance rate relative to the initial response and, without privileging any single query, as the mean pairwise agreement and Fleiss κ across the 4 iterations. Concordance with MDT decisions was assessed using both the initial and modal responses, together with Cohen κ and class-wise F 1 -scores. Results A total of 107 MDT cases were analyzed. Stability differed significantly across models ( P =.01). Gemini 3 Pro showed the lowest discordance rate (mean 12.8%, SD 2.3%) and the highest reference-free agreement (Fleiss κ=0.737), whereas GPT-4o showed the highest discordance rate (mean 30.2%, SD 6.5%) and the lowest agreement (Fleiss κ=0.430). Concordance with MDT decisions ranged from 48.6% to 72.9% using the initial response and from 66.3% to 74.5% using the modal response, and the highest-performing model differed between the 2 definitions. Class-wise F 1 -score was consistently lower for surgery (0.400‐0.520) than for chemotherapy (0.621‐0.836). Complete discordance occurred in 17 of 107 (15.9%) cases and in none of the 31 anatomically unresectable cases (Fisher exact test, P =.003). Recurrent or on-treatment disease (adjusted odds ratio [OR] 5.40, 95% CI 1.65-17.68; P =.005), pancreatic tumor location (adjusted OR 7.37, 95% CI 2.03-26.78; P =.002), and low MDT agreement level (adjusted OR 10.33, 95% CI 1.54-69.38; P =.016) were independently associated with complete discordance. Conclusions LLM-generated treatment recommendations demonstrated moderate alignment with MDT decisions in HPB oncology. Importantly, response stability varied substantially across models, indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools. These findings suggest that LLMs may serve as a reasoning-support layer in MDT-like decision environments, but their response stability must be systematically characterized before clinical integration.

Authors

Publication Details

Journal
Journal of Medical Internet Research
Published
2026-09-17
DOI
https://doi.org/10.2196/101137
Primary Topic
Cholangiocarcinoma and Gallbladder Cancer Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Large Language Models in Multidisciplinary Decision-Making for Hepatopancreatobiliary Oncology: Retrospective Comparative Feasibility Study

Min Je Sung, Incheon Kang, Kwang Hyun Ko, Ho Yeong Lim et al.
Journal of Medical Internet Research
Cholangiocarcinoma and Gallbladder Cancer Studies
article

Large Language Models in Multidisciplinary Decision-Making for Hepatopancreatobiliary Oncology: Retrospective Comparative Feasibility Study

Min Je Sung, Incheon Kang, Kwang Hyun Ko, Ho Yeong Lim, Chansik An, Jung Ho Im, Sung Jun Jo, Hong Jae Chon, Beodeul Kang, Sung Hwan Lee, Jeong‐Sik Yu, Suk Pyo Shin, Eui Hyuk Chong, Jung Sun Kim, Seok Jeong Yang, Sujin Jang, Chang-il Kwon
article en

Abstract

Abstract Background Hepatopancreatobiliary (HPB) malignancies require complex treatment planning that often relies on multidisciplinary team (MDT) discussions. Large language models (LLMs) have recently been explored for clinical decision support, but their performance within real-world multidisciplinary decision environments remains unclear. In particular, the stability of LLM-generated recommendations—that is, whether a model produces the same answer when given the same clinical input—has rarely been examined. Objective This study aimed to evaluate the stability of treatment recommendations generated by contemporary LLMs when identical HPB cases are queried repeatedly, and their concordance with the treatment decisions reached at an institutional MDT conference. Methods This retrospective study included consecutive cases discussed at a single-center HPB MDT conference between September 1, 2024, and August 31, 2025. Standardized clinical case summaries derived from preconference documentation were provided to 4 LLMs (GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5) through their consumer web interfaces. Each model recommended a treatment among predefined MDT treatment options, and identical queries were repeated 4 times in separate sessions. Stability was quantified as the discordance rate relative to the initial response and, without privileging any single query, as the mean pairwise agreement and Fleiss κ across the 4 iterations. Concordance with MDT decisions was assessed using both the initial and modal responses, together with Cohen κ and class-wise F 1 -scores. Results A total of 107 MDT cases were analyzed. Stability differed significantly across models ( P =.01). Gemini 3 Pro showed the lowest discordance rate (mean 12.8%, SD 2.3%) and the highest reference-free agreement (Fleiss κ=0.737), whereas GPT-4o showed the highest discordance rate (mean 30.2%, SD 6.5%) and the lowest agreement (Fleiss κ=0.430). Concordance with MDT decisions ranged from 48.6% to 72.9% using the initial response and from 66.3% to 74.5% using the modal response, and the highest-performing model differed between the 2 definitions. Class-wise F 1 -score was consistently lower for surgery (0.400‐0.520) than for chemotherapy (0.621‐0.836). Complete discordance occurred in 17 of 107 (15.9%) cases and in none of the 31 anatomically unresectable cases (Fisher exact test, P =.003). Recurrent or on-treatment disease (adjusted odds ratio [OR] 5.40, 95% CI 1.65-17.68; P =.005), pancreatic tumor location (adjusted OR 7.37, 95% CI 2.03-26.78; P =.002), and low MDT agreement level (adjusted OR 10.33, 95% CI 1.54-69.38; P =.016) were independently associated with complete discordance. Conclusions LLM-generated treatment recommendations demonstrated moderate alignment with MDT decisions in HPB oncology. Importantly, response stability varied substantially across models, indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools. These findings suggest that LLMs may serve as a reasoning-support layer in MDT-like decision environments, but their response stability must be systematically characterized before clinical integration.

Journal of Medical Internet ResearchVol. 28
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Cholangiocarcinoma and Gallbladder Cancer Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.