Pranay M. Mahendrakar · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22798758
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. An agent that fails a task and writes down what it should have done instead has produced a specific kind of object: a sentence asserting that a named step caused the failure and that a named alternative would not have. This paper takes that object apart. A stored lesson carries three separable claims - that the trajectory failed, that the identified step is why, and that the correction applies to some future task - and each is licensed by a different thing. The first is licensed, in the founding systems, by an exact-match grader, a unit test, or a hand-written stuckness heuristic that fires on action repetition or a step budget. The second is a counterfactual, and no published evaluation of these systems scores it. The third is measured only through end-task success, which cannot separate a lesson being right from a lesson being retrieved. The premise that these systems leave retention unspecified does not survive contact with the papers: Reflexion bounds its store to one to three experiences and evicts by recency, giving context length as the reason; ExpeL gives each insight an importance count that starts at two and deletes it at zero; and the 2026 literature adds decay, budgeted net value, and utility-over-retrievals. What no published rule keys on is whether the second claim held. The strongest evidence that this matters is adversarial to the popular reading and comes from inside the field's own papers: ExpeL's ablation reports that feeding Reflexion-style reflections into its insight extractor drops HotpotQA success from 39.0 to 29.0 against a 28.0 baseline, and attributes the drop to hallucinated reflections; a 2026 framework names the self-confirmation trap, finds that adding self-verification to a single agent slightly lowers performance, and shows that injecting erroneous but internally coherent experience into 10 percent of a memory bank costs 5.3 points of Pass@1. One human audit of stored-lesson correctness was located, covering one domain of one benchmark. This paper states what follows for reading the literature, states flatly what is not known, and names the comparisons that would settle the open part. The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv record or Crossref before inclusion, and every quantitative claim was read back against the cited source's own table or text before it was written down. No experiment was run and no number in this paper was measured by its author; every number is quoted from the paper credited with it. Section 2 states the search procedure and its limits so that the coverage claims in Sections 10 and 13 can be checked and, if wrong, corrected. The author is responsible for the final text and for all claims made in it.
No comments yet — start the discussion below.