Evaluating the use of artificial intelligence (AI) in systematic review abstract screening: a comparative study of freely available AI-aided tools

Abstract Background Manual abstract screening in systematic reviews is a time-consuming and labour-intensive task. With the rise of artificial intelligence (AI), the number of published articles has grown substantially, adding to the workload of review studies that rely on robust and timely evidence synthesis. At the same time, AI-aided screening tools have been developed to accelerate this process. While previous studies have demonstrated the efficiency of such tools, ongoing technological advances necessitate updated evaluations, particularly for tools that are freely available. In review types such as umbrella reviews, where both the topic area and study design are central to eligibility decisions, the performance of AI-aided tools remains underexplored. Methods We conducted a comparative evaluation of six freely available AI-aided abstract screening tools: Rayyan, RobotAnalyst, PICO Portal, Abstrackr, ASReview, and Colandr using a previously completed umbrella review of interdisciplinary urban planning and public health studies. We assessed (1) early recall performance (i.e., identification of included studies within the first 10% and 25% of screening), (2) feature availability and depth, and (3) user experience. This Study Within a Review (SWAR) was registered in the SWAR repository as SWAR 25. Results All evaluated tools supported the review process by facilitating screening and offering features such as prioritization and keyword highlighting. However, none identified more than 50% of the previously included studies within the first 25% of screening. Feature analysis and user feedback suggested that Rayyan and PICO Portal achieved the highest feature analysis scores among the evaluated tools for our interdisciplinary umbrella review context, although limitations were noted in duplicate removal and in recognizing the importance of study design in eligibility decisions. Conclusions Although a growing number of AI-aided abstract screening tools are publicly and freely available, their accuracy, usability, and adaptability to different review designs remain limited. Enhanced support for duplicate detection and integration of study design considerations could improve their utility in umbrella reviews and other complex evidence syntheses. Continued evaluation and user training may support broader adoption across diverse research contexts.

Authors

Institutions

Publication Details

Journal
Systematic Reviews
Published
2026-09-10
DOI
https://doi.org/10.1186/s13643-026-03313-8
Primary Topic
Meta-analysis and systematic reviews
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating the use of artificial intelligence (AI) in systematic review abstract screening: a comparative study of freely available AI-aided tools

Christopher Tate, Selin Akaraci, Sophie Jones, Niamh O’Kane et al.
Systematic Reviews
Meta-analysis and systematic reviews
article

Evaluating the use of artificial intelligence (AI) in systematic review abstract screening: a comparative study of freely available AI-aided tools

Christopher Tate, Selin Akaraci, Sophie Jones, Niamh O’Kane, Maureen Dobbins, Leandro Garcia, Mike Clarke, Ruth F. Hunter
article en

Abstract

Abstract Background Manual abstract screening in systematic reviews is a time-consuming and labour-intensive task. With the rise of artificial intelligence (AI), the number of published articles has grown substantially, adding to the workload of review studies that rely on robust and timely evidence synthesis. At the same time, AI-aided screening tools have been developed to accelerate this process. While previous studies have demonstrated the efficiency of such tools, ongoing technological advances necessitate updated evaluations, particularly for tools that are freely available. In review types such as umbrella reviews, where both the topic area and study design are central to eligibility decisions, the performance of AI-aided tools remains underexplored. Methods We conducted a comparative evaluation of six freely available AI-aided abstract screening tools: Rayyan, RobotAnalyst, PICO Portal, Abstrackr, ASReview, and Colandr using a previously completed umbrella review of interdisciplinary urban planning and public health studies. We assessed (1) early recall performance (i.e., identification of included studies within the first 10% and 25% of screening), (2) feature availability and depth, and (3) user experience. This Study Within a Review (SWAR) was registered in the SWAR repository as SWAR 25. Results All evaluated tools supported the review process by facilitating screening and offering features such as prioritization and keyword highlighting. However, none identified more than 50% of the previously included studies within the first 25% of screening. Feature analysis and user feedback suggested that Rayyan and PICO Portal achieved the highest feature analysis scores among the evaluated tools for our interdisciplinary umbrella review context, although limitations were noted in duplicate removal and in recognizing the importance of study design in eligibility decisions. Conclusions Although a growing number of AI-aided abstract screening tools are publicly and freely available, their accuracy, usability, and adaptability to different review designs remain limited. Enhanced support for duplicate detection and integration of study design considerations could improve their utility in umbrella reviews and other complex evidence syntheses. Continued evaluation and user training may support broader adoption across diverse research contexts.

Systematic Reviews
Queen's University Belfast (GB), McMaster University (CA)
Sustainable cities and communities
Openalex Percentile: Top 9%
Meta-analysis and systematic reviews
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.