Kyle Walker · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23099251
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Many information systems hold far more than any reader can take in at once: search indices, recommendation catalogs, notification streams, security logs, and long-term memory stores for AI assistants. We give a minimal formal model of such systems, built on three assumptions: the store only grows, the attention available to receive surfaced material is bounded, and every surfaced item costs a positive amount of that attention. From these we prove the Hourglass Theorem: the fraction of the store that can be surfaced at any moment is at most $K/N_t$, where $K$ is a fixed capacity and $N_t$ the store size, so it tends to zero and the bound tightens monotonically, for every salience function and every selection rule. We then derive consequences. Fixed relevance thresholds must eventually overflow capacity. If ingestion outpaces capacity, a fixed fraction of the store is never surfaced by any rule. When relevant material is sparse, any filter with a non-vanishing false-positive rate is eventually flooded, so precision must sharpen at rate $O(1/N_t)$, and the context available to the selector must carry at least $k\log_2(N_t/k)$ bits. When relevant material is dense, selection becomes a ranking problem, and store growth weakly benefits an optimal selector. Together: growth rewards selection precision and punishes its absence.
No comments yet — start the discussion below.