Comparing artificial intelligence versus human screening in systematic reviews
Systematic reviews are essential for informing health policy and practice. Artificial intelligence (AI) automates the article screening process and produces time savings, although the performance of AI screening compared to traditional human screening remains uncertain. We undertook this study to compare the performance of two agentic AI tools, namely Loon Lens TM and Catchii, to one another and to humans at the title and abstract screening level. We also compared Loon Lens to humans at the full-text screening level. We developed a de novo research question on the association between any of three ambient air pollutants – carbon monoxide, ozone, nitrogen dioxide – and the onset or worsening of Parkinson’s disease. A health sciences librarian developed the literature search strategy and we proceeded with human and AI screening. Screening results were compared using sensitivity, specificity, positive predictive value, negative predictive value, concordance, kappa, and F1 score. We contrasted these statistics to those obtained by naïve guessing and regressed concordance (agree or disagree with the human reference standard) onto confidence scores provided by Loon Lens, which assigned a confidence level (‘Very High’, ‘High’, Medium’, or ‘Low’) to each of its screening decisions. Human screening was the reference standard against both AI tools; Catchii was the reference standard against Loon Lens. At title and abstract screening, Loon Lens outperformed Catchii when humans were the reference standard (Loon Lens versus human: sensitivity = 0.66, kappa = 0.74, F1 = 0.76; Catchii versus human: sensitivity = 0.49, kappa = 0.46, F1 = 0.50). At full-text screening, most disagreements centered around articles Loon Lens included and humans excluded. At both screening levels, higher confidence scores were associated with lower odds of disagreement between Loon Lens and human screeners. Given the panoply of available AI screening tools and their differential performance, plus the rapidly evolving nature of AI technology, researchers should pilot test their chosen tool at the start of each review. Sensitivity, kappa, and F1 are the optimal performance statistics to employ, especially at title and abstract screening, where the imbalance between proportions of included and excluded citations can inflate concordance and negative predictive value.
Authors
- Mark Oremus (ORCID: https://orcid.org/0000-0001-8190-253X)
- Abel Torres‐Espín (ORCID: https://orcid.org/0000-0002-9787-8738)
- Marco Gorici
Institutions
- University of Waterloo (CA)
Publication Details
- Journal
- Health and Quality of Life Outcomes
- Published
- 2026-10-03
- DOI
- https://doi.org/10.1186/s12955-026-02642-5
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00