ReviewAid: confidence-scored, locally runnable LLM screening and extraction

Systematic reviews are the backbone of evidence synthesis, yet manual full-text screening and data extraction remain severe bottlenecks. While Large Language Models (LLMs) can accelerate these tasks, practical integration is hampered by hallucinations, opaque reasoning, and privacy concerns. This lightning talk introduces ReviewAid (v4.0.0), an open-source, AI-powered tool streamlining PICO-based screening and data extraction while directly addressing these AI challenges.Unlike proprietary solutions, ReviewAid uses a vendor-agnostic architecture supporting cloud providers alongside fully local execution via Ollama, ensuring sensitive data never leaves the researcher's machine. Version 4.0.0 makes reliability central through a four-tier confidence system. Screening evaluates each eligibility criterion separately: the AI reads the paper three independent times, requiring exact supporting quotes for each judgment; agreement yields a high confidence score, disagreement routes the paper to the researcher. For extraction, a deterministic tier checks every AI-extracted value against the paper's own text via exact-match, paraphrase-detection, and negation checks: values found in the text raise the score; hallucinated or contradicted values lower it and are flagged for human verification. When the AI's claimed confidence exceeds what the text supports, the system overrides it downward. Auto-exclusion occurs without human review only when exclusion evidence is unanimous and quote-backed.Operating as a "third reference" layer rather than a human replacement, ReviewAid demonstrates how AI can safely augment the scientific process. This 10-minute talk will outline the screening workflow, confidence tiers in action, proving high-confidence outputs are more reliable than low-confidence ones, and privacy-preserving local deployment, offering a practical framework for integrating AI into evidence synthesis without compromising scientific integrity. - Presented at the FORRT AI in Metascience Online Conference, September 2026.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23018089
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

ReviewAid: confidence-scored, locally runnable LLM screening and extraction

Vihaan Sahu
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
article

ReviewAid: confidence-scored, locally runnable LLM screening and extraction

Vihaan Sahu
article en

Abstract

Systematic reviews are the backbone of evidence synthesis, yet manual full-text screening and data extraction remain severe bottlenecks. While Large Language Models (LLMs) can accelerate these tasks, practical integration is hampered by hallucinations, opaque reasoning, and privacy concerns. This lightning talk introduces ReviewAid (v4.0.0), an open-source, AI-powered tool streamlining PICO-based screening and data extraction while directly addressing these AI challenges.Unlike proprietary solutions, ReviewAid uses a vendor-agnostic architecture supporting cloud providers alongside fully local execution via Ollama, ensuring sensitive data never leaves the researcher's machine. Version 4.0.0 makes reliability central through a four-tier confidence system. Screening evaluates each eligibility criterion separately: the AI reads the paper three independent times, requiring exact supporting quotes for each judgment; agreement yields a high confidence score, disagreement routes the paper to the researcher. For extraction, a deterministic tier checks every AI-extracted value against the paper's own text via exact-match, paraphrase-detection, and negation checks: values found in the text raise the score; hallucinated or contradicted values lower it and are flagged for human verification. When the AI's claimed confidence exceeds what the text supports, the system overrides it downward. Auto-exclusion occurs without human review only when exclusion evidence is unanimous and quote-backed.Operating as a "third reference" layer rather than a human replacement, ReviewAid demonstrates how AI can safely augment the scientific process. This 10-minute talk will outline the screening workflow, confidence tiers in action, proving high-confidence outputs are more reliable than low-confidence ones, and privacy-preserving local deployment, offering a practical framework for integrating AI into evidence synthesis without compromising scientific integrity. - Presented at the FORRT AI in Metascience Online Conference, September 2026.

Zenodo (CERN European Organization for Nuclear Research)
Reduced inequalities
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.