Comparing artificial intelligence versus human screening in systematic reviews

Systematic reviews are essential for informing health policy and practice. Artificial intelligence (AI) automates the article screening process and produces time savings, although the performance of AI screening compared to traditional human screening remains uncertain. We undertook this study to compare the performance of two agentic AI tools, namely Loon Lens TM and Catchii, to one another and to humans at the title and abstract screening level. We also compared Loon Lens to humans at the full-text screening level. We developed a de novo research question on the association between any of three ambient air pollutants – carbon monoxide, ozone, nitrogen dioxide – and the onset or worsening of Parkinson’s disease. A health sciences librarian developed the literature search strategy and we proceeded with human and AI screening. Screening results were compared using sensitivity, specificity, positive predictive value, negative predictive value, concordance, kappa, and F1 score. We contrasted these statistics to those obtained by naïve guessing and regressed concordance (agree or disagree with the human reference standard) onto confidence scores provided by Loon Lens, which assigned a confidence level (‘Very High’, ‘High’, Medium’, or ‘Low’) to each of its screening decisions. Human screening was the reference standard against both AI tools; Catchii was the reference standard against Loon Lens. At title and abstract screening, Loon Lens outperformed Catchii when humans were the reference standard (Loon Lens versus human: sensitivity = 0.66, kappa = 0.74, F1 = 0.76; Catchii versus human: sensitivity = 0.49, kappa = 0.46, F1 = 0.50). At full-text screening, most disagreements centered around articles Loon Lens included and humans excluded. At both screening levels, higher confidence scores were associated with lower odds of disagreement between Loon Lens and human screeners. Given the panoply of available AI screening tools and their differential performance, plus the rapidly evolving nature of AI technology, researchers should pilot test their chosen tool at the start of each review. Sensitivity, kappa, and F1 are the optimal performance statistics to employ, especially at title and abstract screening, where the imbalance between proportions of included and excluded citations can inflate concordance and negative predictive value.

Authors

Institutions

Publication Details

Journal
Health and Quality of Life Outcomes
Published
2026-10-03
DOI
https://doi.org/10.1186/s12955-026-02642-5
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Comparing artificial intelligence versus human screening in systematic reviews

Mark Oremus, Abel Torres‐Espín, Marco Gorici
Health and Quality of Life Outcomes
Artificial Intelligence in Healthcare and Education
article

Comparing artificial intelligence versus human screening in systematic reviews

Mark Oremus, Abel Torres‐Espín, Marco Gorici
article en

Abstract

Systematic reviews are essential for informing health policy and practice. Artificial intelligence (AI) automates the article screening process and produces time savings, although the performance of AI screening compared to traditional human screening remains uncertain. We undertook this study to compare the performance of two agentic AI tools, namely Loon Lens TM and Catchii, to one another and to humans at the title and abstract screening level. We also compared Loon Lens to humans at the full-text screening level. We developed a de novo research question on the association between any of three ambient air pollutants – carbon monoxide, ozone, nitrogen dioxide – and the onset or worsening of Parkinson’s disease. A health sciences librarian developed the literature search strategy and we proceeded with human and AI screening. Screening results were compared using sensitivity, specificity, positive predictive value, negative predictive value, concordance, kappa, and F1 score. We contrasted these statistics to those obtained by naïve guessing and regressed concordance (agree or disagree with the human reference standard) onto confidence scores provided by Loon Lens, which assigned a confidence level (‘Very High’, ‘High’, Medium’, or ‘Low’) to each of its screening decisions. Human screening was the reference standard against both AI tools; Catchii was the reference standard against Loon Lens. At title and abstract screening, Loon Lens outperformed Catchii when humans were the reference standard (Loon Lens versus human: sensitivity = 0.66, kappa = 0.74, F1 = 0.76; Catchii versus human: sensitivity = 0.49, kappa = 0.46, F1 = 0.50). At full-text screening, most disagreements centered around articles Loon Lens included and humans excluded. At both screening levels, higher confidence scores were associated with lower odds of disagreement between Loon Lens and human screeners. Given the panoply of available AI screening tools and their differential performance, plus the rapidly evolving nature of AI technology, researchers should pilot test their chosen tool at the start of each review. Sensitivity, kappa, and F1 are the optimal performance statistics to employ, especially at title and abstract screening, where the imbalance between proportions of included and excluded citations can inflate concordance and negative predictive value.

Health and Quality of Life Outcomes
University of Waterloo (CA)
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.