Dino Vitale, Baixin Guo · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22980060
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
A memory benchmark run against a deployed AI agent with three reader models returned 36%, 40% and 40%, which read as evidence that the reader model barely matters for recall. An external audit of the public artifacts established that the evaluation could not distinguish a retrieval failure from a memory failure from a reader failure. Reconstruction from source showed both of the agent's retrieval paths were inert under the tested condition, so the benchmark had been measuring context-window membership for every reader. With the harness repaired and instrumented, the same three readers score 36%, 48% and 76% on the same questions, and a single-variable comparison moves one reader from 40% to 76%. The paper decomposes the outcome into three gates, documents a case where working retrieval converted honest abstentions into fabrications, reports a four-rater blind judging protocol over 150 predictions, and specifies what an agent memory evaluation must log.
No comments yet — start the discussion below.