Artificial intelligence tools for evidence synthesis in the biomedical field: A living rapid review and evidence map

Abstract Evidence synthesis (ES) is costly and time-intensive. Artificial intelligence (AI) tools are increasingly used to automate or semi-automate ES tasks. We conducted a living rapid review of AI tools that automate or semi-automate ES tasks. We searched for published articles and preprints through MEDLINE, Embase, and the Association for Computing Machinery Digital Library from January 1, 2021 through October 15, 2025, supplemented by hand-searching. We included existing reviews to capture evaluations before 2021. Screening was conducted in duplicate at the title/abstract level, with single full-text screening and data extraction verified by a second reviewer. We included 134 studies. Performance varied by task and evaluation approach. Search automation tools showed poor recall (median 17%), although human-in-the-loop approaches improved performance. Tools for identifying randomized controlled trials demonstrated high median recall (98%) and precision (92%). Semi-automated abstract screening tools using active learning reduced screening burden by approximately 50% while maintaining 95% recall. Large language models demonstrated variable accuracy for abstract screening (median 88% recall), full-text screening (median 96% recall), and data extraction (median 63% accuracy). Risk of bias (RoB) assessment tools showed moderate agreement with human reviewers for the Cochrane RoB tool version 1 (median 71%) but performed less well with other RoB tools. AI tools can meaningfully reduce burden in study design identification and abstract screening but perform inconsistently across other evidence synthesis tasks. Tools for search automation, data extraction, and RoB assessment require further development. No existing tool can currently replace human expertise.

Authors

Institutions

Publication Details

Journal
Research Synthesis Methods
Published
2026-10-08
DOI
https://doi.org/10.1017/rsm.2026.10118
Primary Topic
Meta-analysis and systematic reviews
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Artificial intelligence tools for evidence synthesis in the biomedical field: A living rapid review and evidence map

Ethan M. Balk, Gaelen P. Adam, Htun Ja Mai, Melinda C. Davies et al.
Research Synthesis Methods
Meta-analysis and systematic reviews
article

Artificial intelligence tools for evidence synthesis in the biomedical field: A living rapid review and evidence map

Ethan M. Balk, Gaelen P. Adam, Htun Ja Mai, Melinda C. Davies, Holly R. Wethington, Eduardo Lucia Caputo, Thomas A Trikalinos, Ilya Ivlev, Haley K Holmer, Erin L. Coppola, Edi Kuhn
article en

Abstract

Abstract Evidence synthesis (ES) is costly and time-intensive. Artificial intelligence (AI) tools are increasingly used to automate or semi-automate ES tasks. We conducted a living rapid review of AI tools that automate or semi-automate ES tasks. We searched for published articles and preprints through MEDLINE, Embase, and the Association for Computing Machinery Digital Library from January 1, 2021 through October 15, 2025, supplemented by hand-searching. We included existing reviews to capture evaluations before 2021. Screening was conducted in duplicate at the title/abstract level, with single full-text screening and data extraction verified by a second reviewer. We included 134 studies. Performance varied by task and evaluation approach. Search automation tools showed poor recall (median 17%), although human-in-the-loop approaches improved performance. Tools for identifying randomized controlled trials demonstrated high median recall (98%) and precision (92%). Semi-automated abstract screening tools using active learning reduced screening burden by approximately 50% while maintaining 95% recall. Large language models demonstrated variable accuracy for abstract screening (median 88% recall), full-text screening (median 96% recall), and data extraction (median 63% accuracy). Risk of bias (RoB) assessment tools showed moderate agreement with human reviewers for the Cochrane RoB tool version 1 (median 71%) but performed less well with other RoB tools. AI tools can meaningfully reduce burden in study design identification and abstract screening but perform inconsistently across other evidence synthesis tasks. Tools for search automation, data extraction, and RoB assessment require further development. No existing tool can currently replace human expertise.

Research Synthesis Methods
Kaiser Permanente (US), Brown University (US), Agency for Healthcare Research and Quality (US)
Openalex Percentile: Top 10%
Meta-analysis and systematic reviews
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.