Shashi Kant · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22996163
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This preprint evaluates dense, sparse (BM25), and hybrid retrieval methods fordomain-specific retrieval-augmented generation (RAG) on a synthetic Nginx technicaldocumentation knowledge base (60 documents, 180 test questions across keyword,semantic, and troubleshooting categories). Hybrid retrieval (alpha=0.75) achieves thebest overall Recall@1 (76.1%) and MRR (0.831), though this improvement over dense-onlyretrieval is not statistically significant at n=180 (paired bootstrap and McNemartests). Cross-encoder reranking shows no overall benefit but a category-dependenteffect. Full methodology, results tables, error analysis, and discussion are in theattached PDF. Code, dataset, and reproducible Colab notebooks are available at:https://github.com/shashikantkaushik/rag-advanced-research-codebase
No comments yet — start the discussion below.