Pranay M. Mahendrakar · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23004074
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. Machine learning uses one word for four operations. Catastrophic forgetting is damage that fine-tuning does to earlier capabilities. Transience is the fading of individual training examples during ordinary training. Resets deliberately reinitialise part of a network to restore its ability to learn. Unlearning deliberately removes targeted knowledge. Read side by side, the literatures appear to contradict each other: forgetting is too easy to cause by accident, as when ten fine-tuning examples strip a model's safety behaviour, and too hard to cause on purpose, as when unlearned knowledge returns after a few steps of fine-tuning on unrelated data. This paper argues that the contradiction comes largely from the instruments. Accidental forgetting is usually scored by output accuracy, and deliberate forgetting by adversarial recovery. When accidental forgetting is probed the same way, most of it is recoverable too: in one controlled study task accuracy falls from near 100 percent to about 20 percent while a recovery probe still reaches 96 percent. The paper separates three things a forgetting operation can change: access (whether stored content reaches the output), content (whether it survives cheap recovery) and trainability (how fast the network learns anything new). Gradient-based forgetting, accidental or deliberate, mostly changes access. The operations shown so far to change content are mainly ones that do not carry the trained parameters forward: retraining without the data, distilling into a fresh or noised network, or filtering the data before pretraining. Plasticity resets use the same lever. We ran no experiments. The paper consolidates published measurements, states a decision procedure, and proposes a two-ratio savings protocol that would place any forgetting result on the partition. The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every arXiv citation was machine-verified against its live arXiv Atom API record, and every other citation against its Crossref, PubMed or DOI record, during drafting (title and author list checked against the record returned). Every quantitative claim is taken from the abstract, full text or a table of the source credited with it; full-text numbers were read from the sources' own HTML renderings rather than from summaries. No experiment was run and no number in this paper was measured by its author. Table 1 re-presents numbers published by the cited papers, each named on its row; Figure 1 plots published values from six cited papers with no transformation, and values the sources state as approximate are plotted at the stated approximation and marked as such. The access / content / trainability partition, Table 2, Algorithm 1 and the savings protocol in Section 10 are original conceptual synthesis by the author, not empirical results, and are presented as such. Two classic connectionist references (McCloskey and Cohen 1989; French 1999) were verified as records but their full texts were not available to the verification step; they are cited only for the names they gave the phenomenon.
No comments yet — start the discussion below.