sai sohan merugu · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23096051
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
A language model that keeps learning after deployment must also be able to forget on request. We present DOTS, a retrofit that attaches a 262,144-slot sparse memory to a frozen pretrained transformer and records every learned unit of knowledge, a Dot, in a ledger. Memory addresses are fixed random hypervectors applied to whitened hidden states, so attaching the memory leaves the model's outputs bit-identical and a learned item is never orphaned by address drift. Each Dot writes only value rows that ordinary text rarely reads and that no earlier Dot owns, using a per-Dot optimizer under deterministic kernels. Because every Dot logs the slots it read and wrote, removing a Dot reduces to restoring rows and replaying only the Dots it could have influenced. The result is bitwise identical to a model that never learned the removed Dot. On Qwen2.5-0.5B learning 1,000 TriviaQA facts it did not know, in ten sequential Dots, DOTS recalls 96.1% of the facts (73.5% under a prompt context never seen in training), keeps 86.3% of facts the model already knew, and changes held-out perplexity by 1.6%. The best full fine-tuning run keeps 51.3% of known facts and the best LoRA run 33.0%. We verify bit-exact removal and reset on both CPU and GPU, and we report the main cost: in our setting every removal replays all later Dots.
No comments yet — start the discussion below.