Pavan Kumar Guthikonda · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23096885
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Support systems built on retrieval often store questions whose answers users accepted, linked to the thread that answered them, so that the next similar question can be answered from the resolved case. This is the retain step of case-based reasoning and a change to the retrieval half of a retrieval-augmented generation pipeline. No text is generated. This paper measures what this human-gated loop does on real data, using all twelve CQADupStack forums with 446,641 posts and 23,703 questions, community duplicate links as ground truth, and real repeat questions replayed through three retrievers, bge-small-en-v1.5, all-MiniLM-L6-v2 and BM25. Only user acceptance is simulated. When the stored entry is the question text, write-back with dense retrievers helps while users reject most wrong answers and hurts beyond a crossover acceptance rate that ranges from below 0.10 to about 0.86 and is a property of each configuration, not a general constant. The loss comes from question-to-question similarity, and storing the question together with the shown thread removes most of the loss and most of the gain. A plug-in estimate of the crossover tested out of sample has a mean absolute error of 0.085. A Doc2Query-style filter halves the loss but does less than reopening wrong answers, and serving stored questions as a semantic cache is harmful at loose thresholds for all three retrievers, including BM25. On a second corpus built from MS MARCO query families the same pattern appears with higher crossovers. Version 3 rewrites the related work around the closest prior studies of self-written agent memory, query-based document expansion, similar-question retrieval and semantic caching, and adds the filter and cache baselines and the MS MARCO corpus. Code is at github.com/pavansky/self-updating-retrieval.
No comments yet — start the discussion below.