Reducing hallucination in medical question answering: A comparative study of standard RAG and simplified chained and dynamic retrieval strategies

Hallucinations in medical question-answering systems pose a serious patient safety risk, as LLMs can generate clinically incorrect or unsupported responses. While retrieval-augmented generation (RAG) is widely used to mitigate this problem, no controlled, same-conditions comparison of static, chained, and dynamic retrieval has used automated hallucination scoring. This study empirically compares three RAG strategies under identical experimental conditions: standard RAG, a simplified, computationally approximation of chain-of-retrieval (simplified CoRAG), and a corrected, simplified two-stage approximation of dynamic retrieval (simplified DRAGIN), whose second retrieval stage now excludes passages already retrieved in the first. Neither method reimplements the full original architectures, so results reflect these implementations, not the broader architectures. Evaluation uses 350 clinically relevant questions from the MedQuAD corpus, stratified across four disease categories (oncology, neurological, cardiovascular/pulmonary, and endocrine/digestive/renal), with Meditron-7B as the generation model. Hallucination is measured via an NLI-based entailment pipeline using RoBERTa-large-MNLI, for annotation-free evaluation. Simplified DRAGIN achieved the lowest mean hallucination score ( 0.246 ± 0.258 ), marginally below standard RAG ( 0.252 ± 0.271 )—a 2.5% reduction (Wilcoxon p = 0.003 , Cohen’s d = 0.06 ) significant but practically small. Both standard RAG and simplified DRAGIN clearly outperformed simplified CoRAG ( 0.325 ± 0.312 ; p < 0.0001 , d ≈ 0.36 – 0.40 ). Hallucination was higher for neurological than oncology questions across all methods; CoRAG's disadvantage (0.518 vs. 0.246) was largest, suggesting domain effects partly explain prior findings. We conclude corrected dynamic retrieval offers at best a modest, category-dependent improvement over standard RAG, while chained retrieval remains a clear underperformer. Given corpus overlap issues (since fixed), absence of physician validation, and simplified implementations, these results represent an early-stage signal rather than a conclusive comparison.

Authors

Publication Details

Journal
Journal of Intelligent & Fuzzy Systems
Published
2026-09-28
DOI
https://doi.org/10.1177/18758967261492002
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Reducing hallucination in medical question answering: A comparative study of standard RAG and simplified chained and dynamic retrieval strategies

Shilpa P. Pant, Rutika Arun Khedkar
Journal of Intelligent & Fuzzy Systems
Topic Modeling
article

Reducing hallucination in medical question answering: A comparative study of standard RAG and simplified chained and dynamic retrieval strategies

Shilpa P. Pant, Rutika Arun Khedkar
article en

Abstract

Hallucinations in medical question-answering systems pose a serious patient safety risk, as LLMs can generate clinically incorrect or unsupported responses. While retrieval-augmented generation (RAG) is widely used to mitigate this problem, no controlled, same-conditions comparison of static, chained, and dynamic retrieval has used automated hallucination scoring. This study empirically compares three RAG strategies under identical experimental conditions: standard RAG, a simplified, computationally approximation of chain-of-retrieval (simplified CoRAG), and a corrected, simplified two-stage approximation of dynamic retrieval (simplified DRAGIN), whose second retrieval stage now excludes passages already retrieved in the first. Neither method reimplements the full original architectures, so results reflect these implementations, not the broader architectures. Evaluation uses 350 clinically relevant questions from the MedQuAD corpus, stratified across four disease categories (oncology, neurological, cardiovascular/pulmonary, and endocrine/digestive/renal), with Meditron-7B as the generation model. Hallucination is measured via an NLI-based entailment pipeline using RoBERTa-large-MNLI, for annotation-free evaluation. Simplified DRAGIN achieved the lowest mean hallucination score ( 0.246 ± 0.258 ), marginally below standard RAG ( 0.252 ± 0.271 )—a 2.5% reduction (Wilcoxon p = 0.003 , Cohen’s d = 0.06 ) significant but practically small. Both standard RAG and simplified DRAGIN clearly outperformed simplified CoRAG ( 0.325 ± 0.312 ; p < 0.0001 , d ≈ 0.36 – 0.40 ). Hallucination was higher for neurological than oncology questions across all methods; CoRAG's disadvantage (0.518 vs. 0.246) was largest, suggesting domain effects partly explain prior findings. We conclude corrected dynamic retrieval offers at best a modest, category-dependent improvement over standard RAG, while chained retrieval remains a clear underperformer. Given corpus overlap issues (since fixed), absence of physician validation, and simplified implementations, these results represent an early-stage signal rather than a conclusive comparison.

Journal of Intelligent & Fuzzy Systems
Good health and well-being
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Reducing hallucination in medical question answering: A comparative study of standard RAG and simplified chained and dynamic retrieval strategies — Shilpa P. Pant, Rutika Arun Khedkar · Journal of Intelligent & Fuzzy Systems (2026) | TGRS Research Map | TGRS