toubanjan · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23105432
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Infinite-RoSA (v2.0) Dense Every-Token Retrieval for RWKV-8 via Dynamic Context-Adapted Rank-$K$ Attractors and Speculative Prefetching Abstract (English) This proposal specifies Infinite-RoSA v2.0, an architecture that deepens the mathematical foundation of RWKV-8's RoSA (Rank-One State Architecture) to seamlessly integrate a 10GB+ parametric external memory corpus into local LLMs with strictly O(1) time complexity and zero additional VRAM consumption. Infinite-RoSA v2.0 addresses two critical bottlenecks of previous state-space memory models: multi-memory crosstalk noise during synthesis and geometric misalignment between precomputed static bases and dynamic input contexts. The architecture introduces three mathematical and systems-level innovations: Crosstalk-Free Rank-$K$ State Addition: Eliminates cross-term noise (∑i≠jwiwjkivjT) by applying direct-sum Rank-$K$ outer-product additions onto the state matrix: Satt=∑i∈Bqwi(ki′vi′T) St=diag(e−w)⋅St−1+ktvtT+α⋅Satt Dynamic Context-Adapted Adapter: A lightweight runtime layer conditioned on query qt that aligns precomputed static memory bases (ki,vi) with the current hidden space geometry: ki′=ki+MLPk(ki⊙qt),vi′=vi+MLPv(vi⊙qt) IVF Speculative Cluster Prefetching: Uses probabilistic centroid transition paths P(Cm∣qt,Δqt) to asynchronously stream candidate memory blocks from CPU/NVMe storage into VRAM L1 cache via CUDA Streams, bypassing I/O stalls during decoding. By harmonizing continuous Hopfield Attractor Dynamics ( βt=β0⋅g(qt) ) with RWKV's inherent Time Decay ( e−w ), Infinite-RoSA v2.0 achieves dense every-token retrieval with high factual accuracy, zero crosstalk noise, and robust noise tolerance.Generative AI tools (Gemini 3.1 Pro) were used to brainstorm the hypothesis, refine the concepts, and generate parts of the prototype code.
No comments yet — start the discussion below.