A Comparative Study of ad^Writer and Human Evaluations of Guideline Reporting Quality: Performance of ChatGPT and DeepSeek

ABSTRACT Aim The proliferation of clinical practice guidelines presents significant challenges for ensuring evidence‐based care, since evaluating their trustworthiness with the Reporting Items for practice Guidelines in HealThcare (RIGHT) criteria is both time‐consuming and resource‐intensive. To address this, the Artificial intelligence‐empowered DeVelopment and AssessmeNt of healthCare guidElines and stanDards (ADVANCED) working group developed ad^Writer, an artificial intelligence system, designed to automate guideline reporting quality assessment. This study aimed to determine the agreement and reliability of ad^Writer compared with human evaluations. Methods Guidelines were identified from four published systematic reviews, and through direct communication with the original authors, 60 were ultimately included. These guidelines were evaluated using ad^Writer with two different large language models, each performing three independent assessments. A majority rule was applied to synthesize results. Results from systematic reviews, which served as the reference standard. Agreement, consistency, and stability of the artificial intelligence‐based assessments were examined using Bland–Altman plots, accuracy measures, and intraclass correlation coefficient (ICC), with a focus on identifying patterns of performance across individual RIGHT items. Results Both models showed good agreement with human but had wide limits of agreement. DeepSeek performed better (mean difference = −1.40, standard deviation = 5.00, 95% limits of agreement ([−11.22, 8.42]) and had higher accuracy scores. Both achieved ≥ 75% accuracy on 23 items (65.71%), however, the results varied. ChatGPT showed higher consistency (ICC = 0.93) compared to DeepSeek (ICC = 0.88). Conclusions The ad^Writer system leverages large language models to demonstrate the potential for rapid, accurate, and reproducible guideline quality assessments based on RIGHT. By serving as a supportive tool for manual evaluation, this approach enhances the efficiency of the guideline appraisal process, saving time and human resources for stakeholders.

Authors

Institutions

Publication Details

Journal
Journal of Evidence-Based Medicine
Published
2026-10-06
DOI
https://doi.org/10.1111/jebm.70193
Primary Topic
Clinical practice guidelines implementation
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A Comparative Study of ad^Writer and Human Evaluations of Guideline Reporting Quality: Performance of ChatGPT and DeepSeek

Qianling Shi, Zeming Li, Jie Zhang, Yishan Qin et al.
Journal of Evidence-Based Medicine
Clinical practice guidelines implementation
article

A Comparative Study of ad^Writer and Human Evaluations of Guideline Reporting Quality: Performance of ChatGPT and DeepSeek

Qianling Shi, Zeming Li, Jie Zhang, Yishan Qin, Bingyi Wang, Honghao Lai, Hui Liu, Zhenhua Yang, the ADVANCED Working Group, Xufei Luo, Yaolong Chen
article en

Abstract

ABSTRACT Aim The proliferation of clinical practice guidelines presents significant challenges for ensuring evidence‐based care, since evaluating their trustworthiness with the Reporting Items for practice Guidelines in HealThcare (RIGHT) criteria is both time‐consuming and resource‐intensive. To address this, the Artificial intelligence‐empowered DeVelopment and AssessmeNt of healthCare guidElines and stanDards (ADVANCED) working group developed ad^Writer, an artificial intelligence system, designed to automate guideline reporting quality assessment. This study aimed to determine the agreement and reliability of ad^Writer compared with human evaluations. Methods Guidelines were identified from four published systematic reviews, and through direct communication with the original authors, 60 were ultimately included. These guidelines were evaluated using ad^Writer with two different large language models, each performing three independent assessments. A majority rule was applied to synthesize results. Results from systematic reviews, which served as the reference standard. Agreement, consistency, and stability of the artificial intelligence‐based assessments were examined using Bland–Altman plots, accuracy measures, and intraclass correlation coefficient (ICC), with a focus on identifying patterns of performance across individual RIGHT items. Results Both models showed good agreement with human but had wide limits of agreement. DeepSeek performed better (mean difference = −1.40, standard deviation = 5.00, 95% limits of agreement ([−11.22, 8.42]) and had higher accuracy scores. Both achieved ≥ 75% accuracy on 23 items (65.71%), however, the results varied. ChatGPT showed higher consistency (ICC = 0.93) compared to DeepSeek (ICC = 0.88). Conclusions The ad^Writer system leverages large language models to demonstrate the potential for rapid, accurate, and reproducible guideline quality assessments based on RIGHT. By serving as a supportive tool for manual evaluation, this approach enhances the efficiency of the guideline appraisal process, saving time and human resources for stakeholders.

Journal of Evidence-Based Medicine
Hong Kong Baptist University (HK), Chinese Academy of Medical Sciences & Peking Union Medical College (CN), Lanzhou University (CN)
Openalex Percentile: Top 9%
Clinical practice guidelines implementation
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.