Multi-agent AI for liver disease guideline development and maintenance: a proof-of-concept study

While clinical guidelines remain essential for standardizing medical care, their utility decays as manual updating fails to keep pace with the exponential growth in medical evidence. To address this challenge, we developed a Multi-Agent Retrieval-Augmented Generation (MARG) system dedicated to clinical guideline development and maintenance. We employed a multi-phase evaluation framework to assess MARG’s performance across agent-agent interaction and human-AI collaboration tasks. The system was evaluated against seven 2025 liver disease guidelines, encompassing 137 clinical scenarios. Here we show MARG achieve 77.4% overall concordance with the guidelines, with consistency rates rising from 45.8% under weak consensus to 87.8% at full consensus. Domain-specific performance is 65.5% for diagnosis, 75.0% for prevention, and 81.0% for treatment. MARG also assists in evidence integration, outperforming manual review in recency (6.3% vs. 0.9%), reproducibility (25.5% vs. 7.5%), and quality (24.2% vs. 8.0% Level I evidence). Through four rounds of structured debate, the system resolves 95.8% of recommendation conflicts and achieves strong consensus (>85%) in 74.5% of final outputs. In a small-scale simulation involving 60 scenarios assessed by physicians with 5, 15, and 30 years of experience, MARG is linked to correction of 40-60% of guideline-inconsistent decisions, with the net gain (50%) most apparent in the early-career tier. These proof-of-concept findings suggest that MARG may support clinical guideline development and maintenance. Effective deployment will depend on a tiered human-AI collaboration model built upon curated evidence repositories, machine-interpretable grading protocols, guideline-specific benchmarks, and independent ethical governance. Clinical practice guidelines help doctors make consistent, evidence-based decisions. However, manual updating often falls behind new research, risking outdated recommendations. To address this, we developed an AI system called MARG to review and synthesize new evidence for guideline updates. Tested against seven liver disease guidelines, MARG accurately reproduced most recommendations, identified higher-quality evidence, and resolved disagreements through structured AI debate. In a small simulation with clinicians of varying experience, MARG resolve guideline inconsistencies, especially for early-career doctors. While promising, the system is not meant to replace human expertise. Future work should focus on real-world testing to ensure safety, transparency, and ethical governance before any clinical use. Here the authors develop multi-agent AI that acts like a digital committee to review evidence and flag potential guideline updates. In liver disease benchmarks, it replicated expert consensus and cut early-career clinicians’ guideline errors by half.

Authors

Institutions

Publication Details

Journal
Communications Medicine
Published
2026-10-07
DOI
https://doi.org/10.1038/s43856-026-01961-4
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Multi-agent AI for liver disease guideline development and maintenance: a proof-of-concept study

Xin Wen, Yafan Wang, Huiying Rao, 孔祥沙 et al.
Communications Medicine
Artificial Intelligence in Healthcare and Education
article

Multi-agent AI for liver disease guideline development and maintenance: a proof-of-concept study

Xin Wen, Yafan Wang, Huiying Rao, 孔祥沙, Tianyi Xue, Rui Huang, 马慧, Wei Wang, Zixing Wang, Ran Fei, Wei Han, Ke Yin, Cuiling Ma, Yaoda Hu, Dongbo Chen, Rui Shen, Shaoping She, Suzhen Jiang, Bo Feng, Xiaoxiao Wang, Feng Liu
article en

Abstract

While clinical guidelines remain essential for standardizing medical care, their utility decays as manual updating fails to keep pace with the exponential growth in medical evidence. To address this challenge, we developed a Multi-Agent Retrieval-Augmented Generation (MARG) system dedicated to clinical guideline development and maintenance. We employed a multi-phase evaluation framework to assess MARG’s performance across agent-agent interaction and human-AI collaboration tasks. The system was evaluated against seven 2025 liver disease guidelines, encompassing 137 clinical scenarios. Here we show MARG achieve 77.4% overall concordance with the guidelines, with consistency rates rising from 45.8% under weak consensus to 87.8% at full consensus. Domain-specific performance is 65.5% for diagnosis, 75.0% for prevention, and 81.0% for treatment. MARG also assists in evidence integration, outperforming manual review in recency (6.3% vs. 0.9%), reproducibility (25.5% vs. 7.5%), and quality (24.2% vs. 8.0% Level I evidence). Through four rounds of structured debate, the system resolves 95.8% of recommendation conflicts and achieves strong consensus (>85%) in 74.5% of final outputs. In a small-scale simulation involving 60 scenarios assessed by physicians with 5, 15, and 30 years of experience, MARG is linked to correction of 40-60% of guideline-inconsistent decisions, with the net gain (50%) most apparent in the early-career tier. These proof-of-concept findings suggest that MARG may support clinical guideline development and maintenance. Effective deployment will depend on a tiered human-AI collaboration model built upon curated evidence repositories, machine-interpretable grading protocols, guideline-specific benchmarks, and independent ethical governance. Clinical practice guidelines help doctors make consistent, evidence-based decisions. However, manual updating often falls behind new research, risking outdated recommendations. To address this, we developed an AI system called MARG to review and synthesize new evidence for guideline updates. Tested against seven liver disease guidelines, MARG accurately reproduced most recommendations, identified higher-quality evidence, and resolved disagreements through structured AI debate. In a small simulation with clinicians of varying experience, MARG resolve guideline inconsistencies, especially for early-career doctors. While promising, the system is not meant to replace human expertise. Future work should focus on real-world testing to ensure safety, transparency, and ethical governance before any clinical use. Here the authors develop multi-agent AI that acts like a digital committee to review evidence and flag potential guideline updates. In liver disease benchmarks, it replicated expert consensus and cut early-career clinicians’ guideline errors by half.

Communications Medicine
Capital Medical University (CN), Chinese Academy of Medical Sciences & Peking Union Medical College (CN), Peking University (CN), Beijing Jishuitan Hospital (CN), Beijing Geriatric Hospital (CN), Peking University People's Hospital (CN), Institute of Basic Medical Sciences of the Chinese Academy of Medical Sciences
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.