Multicenter Evaluation of Large Language Models Versus Hepatologists for Prognostic Prediction in Drug‐Induced Liver Injury

BACKGROUND & AIMS: Drug-induced liver injury (DILI) would progress to chronicity or death. Large language models (LLMs) may enhance clinical decision-making, yet their utility relative to physicians in DILI remains unclear. Therefore, we evaluated their performance in predicting DILI outcomes. METHODS: We enrolled 943 DILI patients from three centers as internal and external cohorts. Based on 12-month follow-up, outcomes were classified as recovery, 6/12-month chronicity, and death. LLMs (Gemini-2.5 Pro, GPT-5.1, DeepSeek-3.2), hepatologists (Junior, middle, senior), and models (Hy's Law, nHy's Law, MELD Score) estimated probabilities of outcomes. LLM-Rules (VOTE, OR, AND) were applied to enhance stability. Model performance was assessed. RESULTS: For 6-month chronicity, the senior achieved highest AUROC (0.61) with an accuracy of 70%. Gemini-2.5 Pro and GPT-5.1 yielded AUROCs of 0.60 and 0.59, respectively, outperforming junior and middle hepatologists. Gemini-2.5 Pro demonstrated strongest agreement with senior (κ = 0.43). LLMs all exhibited lower accuracy and specificity than hepatologists. A similar result was observed in 12-month chronicity. For overall mortality, the senior achieved highest AUROC (0.87) with an accuracy of 83%. Gemini-2.5 Pro and GPT-5.1 achieved AUROCs of 0.86, outperforming junior hepatologist, Hy's Law, and nHy's Law. GPT-5.1 achieved strongest agreement with senior (κ = 0.25). LLM-Rules demonstrated stability for predicting outcomes across cohorts. OR and AND rules improved sensitivity and specificity, respectively. CONCLUSIONS: GPT-5.1 and Gemini-2.5 Pro showed AUROCs approaching senior hepatologists for DILI outcomes with limited accuracy and specificity. LLM-Rules demonstrated stable performance across cohorts with improved sensitivity or specificity, supporting the potential of multi-LLM approaches as clinician-supervised complementary tools.

Authors

Institutions

Publication Details

Journal
Liver International
Published
2026-09-18
DOI
https://doi.org/10.1111/liv.70879
Primary Topic
Drug-Induced Hepatotoxicity and Protection
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multicenter Evaluation of Large Language Models Versus Hepatologists for Prognostic Prediction in Drug‐Induced Liver Injury

Rongtao Lai, Yanan Du, Qing Xie, Haoshuang Fu et al.
Liver International
Drug-Induced Hepatotoxicity and Protection
article

Multicenter Evaluation of Large Language Models Versus Hepatologists for Prognostic Prediction in Drug‐Induced Liver Injury

Rongtao Lai, Yanan Du, Qing Xie, Haoshuang Fu, Tianhui Zhou, Gangde Zhao, Shuying Song, Yuelin Xiao, Xinya Zang, Ruidong Mo, Yan Zhuang
article en

Abstract

BACKGROUND & AIMS: Drug-induced liver injury (DILI) would progress to chronicity or death. Large language models (LLMs) may enhance clinical decision-making, yet their utility relative to physicians in DILI remains unclear. Therefore, we evaluated their performance in predicting DILI outcomes. METHODS: We enrolled 943 DILI patients from three centers as internal and external cohorts. Based on 12-month follow-up, outcomes were classified as recovery, 6/12-month chronicity, and death. LLMs (Gemini-2.5 Pro, GPT-5.1, DeepSeek-3.2), hepatologists (Junior, middle, senior), and models (Hy's Law, nHy's Law, MELD Score) estimated probabilities of outcomes. LLM-Rules (VOTE, OR, AND) were applied to enhance stability. Model performance was assessed. RESULTS: For 6-month chronicity, the senior achieved highest AUROC (0.61) with an accuracy of 70%. Gemini-2.5 Pro and GPT-5.1 yielded AUROCs of 0.60 and 0.59, respectively, outperforming junior and middle hepatologists. Gemini-2.5 Pro demonstrated strongest agreement with senior (κ = 0.43). LLMs all exhibited lower accuracy and specificity than hepatologists. A similar result was observed in 12-month chronicity. For overall mortality, the senior achieved highest AUROC (0.87) with an accuracy of 83%. Gemini-2.5 Pro and GPT-5.1 achieved AUROCs of 0.86, outperforming junior hepatologist, Hy's Law, and nHy's Law. GPT-5.1 achieved strongest agreement with senior (κ = 0.25). LLM-Rules demonstrated stability for predicting outcomes across cohorts. OR and AND rules improved sensitivity and specificity, respectively. CONCLUSIONS: GPT-5.1 and Gemini-2.5 Pro showed AUROCs approaching senior hepatologists for DILI outcomes with limited accuracy and specificity. LLM-Rules demonstrated stable performance across cohorts with improved sensitivity or specificity, supporting the potential of multi-LLM approaches as clinician-supervised complementary tools.

Liver InternationalVol. 46(10)
Shanghai Jiao Tong University (CN), Ruijin Hospital (CN), Wuxi People's Hospital (CN)
National Natural Science Foundation of China
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Drug-Induced Hepatotoxicity and Protection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.