Reducing hallucination in medical question answering: A comparative study of standard RAG and simplified chained and dynamic retrieval strategies
Hallucinations in medical question-answering systems pose a serious patient safety risk, as LLMs can generate clinically incorrect or unsupported responses. While retrieval-augmented generation (RAG) is widely used to mitigate this problem, no controlled, same-conditions comparison of static, chained, and dynamic retrieval has used automated hallucination scoring. This study empirically compares three RAG strategies under identical experimental conditions: standard RAG, a simplified, computationally approximation of chain-of-retrieval (simplified CoRAG), and a corrected, simplified two-stage approximation of dynamic retrieval (simplified DRAGIN), whose second retrieval stage now excludes passages already retrieved in the first. Neither method reimplements the full original architectures, so results reflect these implementations, not the broader architectures. Evaluation uses 350 clinically relevant questions from the MedQuAD corpus, stratified across four disease categories (oncology, neurological, cardiovascular/pulmonary, and endocrine/digestive/renal), with Meditron-7B as the generation model. Hallucination is measured via an NLI-based entailment pipeline using RoBERTa-large-MNLI, for annotation-free evaluation. Simplified DRAGIN achieved the lowest mean hallucination score ( 0.246 ± 0.258 ), marginally below standard RAG ( 0.252 ± 0.271 )—a 2.5% reduction (Wilcoxon p = 0.003 , Cohen’s d = 0.06 ) significant but practically small. Both standard RAG and simplified DRAGIN clearly outperformed simplified CoRAG ( 0.325 ± 0.312 ; p < 0.0001 , d ≈ 0.36 – 0.40 ). Hallucination was higher for neurological than oncology questions across all methods; CoRAG's disadvantage (0.518 vs. 0.246) was largest, suggesting domain effects partly explain prior findings. We conclude corrected dynamic retrieval offers at best a modest, category-dependent improvement over standard RAG, while chained retrieval remains a clear underperformer. Given corpus overlap issues (since fixed), absence of physician validation, and simplified implementations, these results represent an early-stage signal rather than a conclusive comparison.
Authors
- Shilpa P. Pant (ORCID: https://orcid.org/0000-0002-6709-5548)
- Rutika Arun Khedkar
Publication Details
- Journal
- Journal of Intelligent & Fuzzy Systems
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1177/18758967261492002
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00