SPIRIT-CONSORT-ELM: element-level annotated dataset and large language model approach for assessing randomized controlled trial reporting

Abstract Randomized controlled trials (RCTs) are central to assessing the benefits and harms of interventions, but incomplete reporting undermines their verifiability and usefulness. Although SPIRIT and CONSORT reporting guidelines promote complete reporting of RCT protocols and results publications, many RCTs remain incompletely reported. Automated manuscript checking could help improve reporting completeness before publication. We previously developed SPIRIT-CONSORT-TM, a corpus of 200 articles (100 protocol-results publication pairs) annotated with 83 checklist items from SPIRIT 2013 and CONSORT 2010, and trained models for item-level assessment. However, checklist items may comprise multiple constituent elements, which prior work did not capture or evaluate. Here, we extend the corpus with element-level annotations (SPIRIT-CONSORT-ELM) and formulate assessment as a machine reading comprehension task operationalized through 119 questions targeting specific reporting elements. Two annotators independently assessed 50 articles (25 pairs), with discrepancies resolved through discussion; one annotator assessed the remaining 150 articles. We then developed an automated pipeline combining PubMedBERT-based evidence retrieval with GPT-5-based question answering. Inter-annotator agreement was high (Gwet’s AC1: 0.782), and the pipeline achieved high performance (F1: 0.822, Gwet’s AC1: 0.796). Component analyses demonstrated the importance of evidence retrieval quality and modest benefits from illustrative in-context examples. SPIRIT-CONSORT-ELM provides a benchmark for fine-grained assessment of RCT reporting completeness, while the automated pipeline establishes a robust baseline and shows potential for supporting authors, reviewers, and editors.

Authors

Publication Details

Journal
npj Digital Medicine
Published
2026-10-06
DOI
https://doi.org/10.1038/s41746-026-03318-6
Primary Topic
Meta-analysis and systematic reviews
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

SPIRIT-CONSORT-ELM: element-level annotated dataset and large language model approach for assessing randomized controlled trial reporting

Andrew William Brown, Halil Kilicoglu, Xiangji Ying, Evan Mayo‐Wilson et al.
npj Digital Medicine
Meta-analysis and systematic reviews
article

SPIRIT-CONSORT-ELM: element-level annotated dataset and large language model approach for assessing randomized controlled trial reporting

Andrew William Brown, Halil Kilicoglu, Xiangji Ying, Evan Mayo‐Wilson, Colby J. Vorland, Joseph T. Menke, Lan Jiang, Mengfei Lan, Wenxuan Song
article en

Abstract

Abstract Randomized controlled trials (RCTs) are central to assessing the benefits and harms of interventions, but incomplete reporting undermines their verifiability and usefulness. Although SPIRIT and CONSORT reporting guidelines promote complete reporting of RCT protocols and results publications, many RCTs remain incompletely reported. Automated manuscript checking could help improve reporting completeness before publication. We previously developed SPIRIT-CONSORT-TM, a corpus of 200 articles (100 protocol-results publication pairs) annotated with 83 checklist items from SPIRIT 2013 and CONSORT 2010, and trained models for item-level assessment. However, checklist items may comprise multiple constituent elements, which prior work did not capture or evaluate. Here, we extend the corpus with element-level annotations (SPIRIT-CONSORT-ELM) and formulate assessment as a machine reading comprehension task operationalized through 119 questions targeting specific reporting elements. Two annotators independently assessed 50 articles (25 pairs), with discrepancies resolved through discussion; one annotator assessed the remaining 150 articles. We then developed an automated pipeline combining PubMedBERT-based evidence retrieval with GPT-5-based question answering. Inter-annotator agreement was high (Gwet’s AC1: 0.782), and the pipeline achieved high performance (F1: 0.822, Gwet’s AC1: 0.796). Component analyses demonstrated the importance of evidence retrieval quality and modest benefits from illustrative in-context examples. SPIRIT-CONSORT-ELM provides a benchmark for fine-grained assessment of RCT reporting completeness, while the automated pipeline establishes a robust baseline and shows potential for supporting authors, reviewers, and editors.

npj Digital Medicine
Openalex Percentile: Top 9%
Meta-analysis and systematic reviews
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.