Ahmed N. Farrag, Ahmed S. A. Soliman, Chidimma Doris Azubuike, Kimia Zandbiglari, Surya Yadavilli, Lauren Adkins, Fatemeh Mehrabi, Larisa Cavallari, Amie J. Goodin, Masoud Rouhizadeh · npj Digital Medicine 2026 · 2026
DOI: 10.1038/s41746-026-03277-y
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Retrieval-augmented and LLM-based (RAG-LLM) evidence-search tools, including Consensus, Ai2 Paper Finder, ChatGPT, Gemini, and Claude, are increasingly used by clinicians and researchers. Whether a realistic query reliably surfaces relevant evidence or leaves systematic gaps that could shape clinical evidence and research synthesis remains unclear. Using a prospectively assembled, non-public gold-standard corpus to avoid benchmark contamination, we assessed five platforms across 15 query formulations. Primary outcomes were formulation-level recall (evidence retrieved per query formulation) and single-query zero-retrieval probability (formulations returning no relevant evidence from a domain); pooled platform recall (evidence retrieved at least once across all formulations) was a secondary capacity benchmark. Median formulation-level recall ranged from 7.2% to 42.2%, while pooled platform recall ranged from 45.8% to 72.3%. For the largest evidence category, single-query zero-retrieval probability ranged from 47% to 80% across platforms; one platform showed a marked pre-2016 evidence gap; and 12.0% of evidence was never retrieved by any platform, with never-retrieval significantly higher for conference proceedings than journal articles (38.9% vs 4.6%; p < 0.001). Evidence gaps varied by platform, evidence category, publication year, and venue type, highlighting potential retrieval bias and visibility blind spots, and supporting domain-specific evaluation before RAG-LLM outputs are used in clinical or research workflows.
No comments yet — start the discussion below.