Rutika Arun Khedkar, Shilpa P. Pant · Journal of Intelligent & Fuzzy Systems 2026 · 2026
DOI: 10.1177/18758967261492002
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Hallucinations in medical question-answering systems pose a serious patient safety risk, as LLMs can generate clinically incorrect or unsupported responses. While retrieval-augmented generation (RAG) is widely used to mitigate this problem, no controlled, same-conditions comparison of static, chained, and dynamic retrieval has used automated hallucination scoring. This study empirically compares three RAG strategies under identical experimental conditions: standard RAG, a simplified, computationally approximation of chain-of-retrieval (simplified CoRAG), and a corrected, simplified two-stage approximation of dynamic retrieval (simplified DRAGIN), whose second retrieval stage now excludes passages already retrieved in the first. Neither method reimplements the full original architectures, so results reflect these implementations, not the broader architectures. Evaluation uses 350 clinically relevant questions from the MedQuAD corpus, stratified across four disease categories (oncology, neurological, cardiovascular/pulmonary, and endocrine/digestive/renal), with Meditron-7B as the generation model. Hallucination is measured via an NLI-based entailment pipeline using RoBERTa-large-MNLI, for annotation-free evaluation. Simplified DRAGIN achieved the lowest mean hallucination score ( 0.246 ± 0.258 ), marginally below standard RAG ( 0.252 ± 0.271 )—a 2.5% reduction (Wilcoxon p = 0.003 , Cohen’s d = 0.06 ) significant but practically small. Both standard RAG and simplified DRAGIN clearly outperformed simplified CoRAG ( 0.325 ± 0.312 ; p < 0.0001 , d ≈ 0.36 – 0.40 ). Hallucination was higher for neurological than oncology questions across all methods; CoRAG's disadvantage (0.518 vs. 0.246) was largest, suggesting domain effects partly explain prior findings. We conclude corrected dynamic retrieval offers at best a modest, category-dependent improvement over standard RAG, while chained retrieval remains a clear underperformer. Given corpus overlap issues (since fixed), absence of physician validation, and simplified implementations, these results represent an early-stage signal rather than a conclusive comparison.
No comments yet — start the discussion below.