Rishabh Sharma, Rishika Lall · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22985242
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound −3.0 points against a −5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061× less to write, and the result holds with a second answer model. Within this study, reranking's gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound −2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released. Code: https://github.com/ris3abh/Engram (tag paper-v3-preprint-r2) Pre-registered plan: 10.5281/zenodo.22970745 Amendment and pre-written outcome paragraphs: 10.5281/zenodo.22977848 Earlier preprint: 10.5281/zenodo.22941757 Data: https://huggingface.co/datasets/ris3abh-11/engram-eval
No comments yet — start the discussion below.