Yao Guohui, Ying Tang, Haiming Wu, Jingxu Cao, Yi Sui, Dawei Song, Songkun Ji · Information Processing & Management 2026 · 2026
DOI: 10.1016/j.ipm.2026.105186
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Retrieval-Augmented Generation (RAG) mitigates hallucinations in Large Language Models (LLMs) by grounding generation in external documents, but retrieving more documents does not always improve answer quality. Additional documents may introduce redundancy or conflicting evidence, leading to non-monotonic performance gains. To address this issue, we propose U-RAG, a Utility-Aware RAG framework that explicitly models the marginal gain of each retrieved document. U-RAG introduces a Utility Scoring Model that estimates the expected improvement in generation quality when a candidate document is added to the current context. Based on this utility estimation, U-RAG adaptively selects useful documents and determines when retrieval should stop, resulting in an average retrieval length of 2.42 under a maximum budget of three. This design enables the model to estimate a local stopping boundary between beneficial and non-beneficial one-step retrieval actions, rather than assuming that more documents are always better. We evaluate U-RAG on five open-domain question answering datasets: 2WikiMultiHopQA, HotpotQA, NaturalQA, TriviaQA, and WebQuestions. Experimental results show that U-RAG achieves the best average performance across these datasets, yielding an average relative improvement of approximately 8.5% in the combined EM/F1 metric over standard RAG. Notably, U-RAG obtains the strongest results among the compared methods on 2WikiMultiHopQA, TriviaQA, and WebQuestions, while remaining competitive with strong retrieval baselines on NaturalQA and HotpotQA.
No comments yet — start the discussion below.