Helong Hu, Hongdan Pan · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23092649
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The quadratic cost of self-attention limits long-context language models, and sparse or linearized alternatives have been judged almost solely by perplexity, leaving open whether they retain the functional abilities of full attention. We show that this metric is insufficient. At 16k context, standard full attention itself loses near-field retrieval, while windowed attention---despite the best validation loss in its class---scores 0\% on a passkey probe and remains at 0\% after 500 steps of targeted fine-tuning, its answer loss plateauing at chance level; the deficit is structural, not a training accident. Inspired by the Complementary Learning Systems account of biological memory, we introduce Episodic--Semantic Attention (ESA): each token attends exactly within its block and to mean-pooled summaries of all strictly previous blocks, both score types competing in a single joint softmax. ESA adds no parameters and uses 1.8\% of full attention's scores. At 16k it is the only architecture tested that combines better-than-baseline validation loss with 65\% retrieval where every baseline scores 0\%; the perplexity advantage is directionally consistent across two seeds but not yet statistically significant at $n=2$. Semantic capacity saturates at $\approx$32 slots, and hierarchical routing emerges with no architectural prior. Perplexity alone therefore overstates windowed methods, and functional retrieval probes should join standard long-context evaluation.
No comments yet — start the discussion below.