Blind spots in AI-assisted healthcare evidence search: multiplatform evaluation of clinical retrieval gaps and risk-of-bias

Abstract Retrieval-augmented and LLM-based (RAG-LLM) evidence-search tools, including Consensus, Ai2 Paper Finder, ChatGPT, Gemini, and Claude, are increasingly used by clinicians and researchers. Whether a realistic query reliably surfaces relevant evidence or leaves systematic gaps that could shape clinical evidence and research synthesis remains unclear. Using a prospectively assembled, non-public gold-standard corpus to avoid benchmark contamination, we assessed five platforms across 15 query formulations. Primary outcomes were formulation-level recall (evidence retrieved per query formulation) and single-query zero-retrieval probability (formulations returning no relevant evidence from a domain); pooled platform recall (evidence retrieved at least once across all formulations) was a secondary capacity benchmark. Median formulation-level recall ranged from 7.2% to 42.2%, while pooled platform recall ranged from 45.8% to 72.3%. For the largest evidence category, single-query zero-retrieval probability ranged from 47% to 80% across platforms; one platform showed a marked pre-2016 evidence gap; and 12.0% of evidence was never retrieved by any platform, with never-retrieval significantly higher for conference proceedings than journal articles (38.9% vs 4.6%; p < 0.001). Evidence gaps varied by platform, evidence category, publication year, and venue type, highlighting potential retrieval bias and visibility blind spots, and supporting domain-specific evaluation before RAG-LLM outputs are used in clinical or research workflows.

Authors

Publication Details

Journal
npj Digital Medicine
Published
2026-09-29
DOI
https://doi.org/10.1038/s41746-026-03277-y
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Blind spots in AI-assisted healthcare evidence search: multiplatform evaluation of clinical retrieval gaps and risk-of-bias

Fatemeh Mehrabi, Amie J. Goodin, Masoud Rouhizadeh, Lauren Adkins et al.
npj Digital Medicine
Artificial Intelligence in Healthcare and Education
article

Blind spots in AI-assisted healthcare evidence search: multiplatform evaluation of clinical retrieval gaps and risk-of-bias

Fatemeh Mehrabi, Amie J. Goodin, Masoud Rouhizadeh, Lauren Adkins, Ahmed S. A. Soliman, Ahmed N. Farrag, Kimia Zandbiglari, Chidimma Doris Azubuike, Surya Yadavilli, Larisa Cavallari
article en

Abstract

Abstract Retrieval-augmented and LLM-based (RAG-LLM) evidence-search tools, including Consensus, Ai2 Paper Finder, ChatGPT, Gemini, and Claude, are increasingly used by clinicians and researchers. Whether a realistic query reliably surfaces relevant evidence or leaves systematic gaps that could shape clinical evidence and research synthesis remains unclear. Using a prospectively assembled, non-public gold-standard corpus to avoid benchmark contamination, we assessed five platforms across 15 query formulations. Primary outcomes were formulation-level recall (evidence retrieved per query formulation) and single-query zero-retrieval probability (formulations returning no relevant evidence from a domain); pooled platform recall (evidence retrieved at least once across all formulations) was a secondary capacity benchmark. Median formulation-level recall ranged from 7.2% to 42.2%, while pooled platform recall ranged from 45.8% to 72.3%. For the largest evidence category, single-query zero-retrieval probability ranged from 47% to 80% across platforms; one platform showed a marked pre-2016 evidence gap; and 12.0% of evidence was never retrieved by any platform, with never-retrieval significantly higher for conference proceedings than journal articles (38.9% vs 4.6%; p < 0.001). Evidence gaps varied by platform, evidence category, publication year, and venue type, highlighting potential retrieval bias and visibility blind spots, and supporting domain-specific evaluation before RAG-LLM outputs are used in clinical or research workflows.

npj Digital Medicine
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.