Sharvesh Sathish Kumar · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23049830
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Transformer-based Large Language Models (LLMs) suffer from memory bottlenecks during long-context inference due to the $O(N)$ spatial footprint of the Key-Value (KV) cache. We propose Continuous Implicit Manifold Attention (CIMA), a novel caching architecture that approximates discrete sequence keys as continuous implicit neural fields using Sinusoidal Representation Networks (SIRENs). To prevent information loss at high-entropy token boundaries, CIMA incorporates a Sparse Residual Anchor Map ($\mathcal{A}$) that selectively retains exact key representations for the top 2% highest reconstruction error states. Evaluated across sequence lengths up to 32,000 tokens, CIMA achieves up to 97.4% VRAM footprint reduction compared to standard FP16 KV caching while maintaining 100% retrieval accuracy on long-context Needle-In-A-Haystack (NIAH) benchmarks.
No comments yet — start the discussion below.