A Comparative Study of ad^Writer and Human Evaluations of Guideline Reporting Quality: Performance of ChatGPT and DeepSeek
ABSTRACT Aim The proliferation of clinical practice guidelines presents significant challenges for ensuring evidence‐based care, since evaluating their trustworthiness with the Reporting Items for practice Guidelines in HealThcare (RIGHT) criteria is both time‐consuming and resource‐intensive. To address this, the Artificial intelligence‐empowered DeVelopment and AssessmeNt of healthCare guidElines and stanDards (ADVANCED) working group developed ad^Writer, an artificial intelligence system, designed to automate guideline reporting quality assessment. This study aimed to determine the agreement and reliability of ad^Writer compared with human evaluations. Methods Guidelines were identified from four published systematic reviews, and through direct communication with the original authors, 60 were ultimately included. These guidelines were evaluated using ad^Writer with two different large language models, each performing three independent assessments. A majority rule was applied to synthesize results. Results from systematic reviews, which served as the reference standard. Agreement, consistency, and stability of the artificial intelligence‐based assessments were examined using Bland–Altman plots, accuracy measures, and intraclass correlation coefficient (ICC), with a focus on identifying patterns of performance across individual RIGHT items. Results Both models showed good agreement with human but had wide limits of agreement. DeepSeek performed better (mean difference = −1.40, standard deviation = 5.00, 95% limits of agreement ([−11.22, 8.42]) and had higher accuracy scores. Both achieved ≥ 75% accuracy on 23 items (65.71%), however, the results varied. ChatGPT showed higher consistency (ICC = 0.93) compared to DeepSeek (ICC = 0.88). Conclusions The ad^Writer system leverages large language models to demonstrate the potential for rapid, accurate, and reproducible guideline quality assessments based on RIGHT. By serving as a supportive tool for manual evaluation, this approach enhances the efficiency of the guideline appraisal process, saving time and human resources for stakeholders.
Authors
- Qianling Shi (ORCID: https://orcid.org/0000-0002-0932-1543)
- Zeming Li (ORCID: https://orcid.org/0009-0002-0895-4354)
- Jie Zhang (ORCID: https://orcid.org/0000-0002-2497-725X)
- Yishan Qin (ORCID: https://orcid.org/0000-0003-2931-8867)
- Bingyi Wang (ORCID: https://orcid.org/0009-0004-8834-893X)
- Honghao Lai
- Hui Liu
- Zhenhua Yang
- the ADVANCED Working Group
- Xufei Luo
- Yaolong Chen
Institutions
- Hong Kong Baptist University (HK)
- Chinese Academy of Medical Sciences & Peking Union Medical College (CN)
- Lanzhou University (CN)
Publication Details
- Journal
- Journal of Evidence-Based Medicine
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1111/jebm.70193
- Primary Topic
- Clinical practice guidelines implementation
- Type
- article
- Field-Weighted Citation Impact
- 0.00