Purvish Haresh Sharma · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23101884
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large Language Models (LLMs) can give fluent and useful answers, but they can alsoproduce information that is false or not supported by evidence. This problem is often calledhallucination. In this paper, we study how the quality of retrieved evidence is associatedwith differences in the answers given by an LLM.The experiment used 13 questions in English, Hindi, and Gujarati. The questions weretested under three conditions: No Retrieval, High-quality Retrieval, and Low-quality Re-trieval. In total, the experiment produced 117 responses. The responses were evaluated us-ing factual correctness, evidence support, appropriate abstention, over-cautious behaviour,unsupported inference, and non-evaluable cases.In the No Retrieval condition, all 39 responses were fully correct. In the High-qualityRetrieval condition, 34 of 36 evaluable responses were fully correct. In the Low-qualityRetrieval condition, only 3 of 39 responses were fully correct, but 36 responses showedappropriate abstention. This means that the model often responded that the evidence wasnot sufficient instead of making an unsupported claim.The results suggest that retrieval does not automatically improve the reliability of anLLM. The quality, relevance, and completeness of the evidence are also important. Thestudy also shows that factual correctness, evidence faithfulness, and appropriate uncertaintyshould be evaluated separately.
No comments yet — start the discussion below.