Adam Allcock · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23121596
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Blindboard tests whether a language model can track facts as they change. The model plays on a keyboard whose layout is hidden. Each turn it presses some keys and is told which symbols came back, but not which key produced which. Between turns some keys trade symbols, and the model is told which keys moved but not where their symbols went. When ready, or at a turn limit, it writes down the whole layout and is scored on the fraction of keys it gets right. Harder versions, called tiers, add keys, keep the board moving, and rename keys mid-game. A reference solver plays every tier first, and a replay tracker shows which missed keys the model's observations had already pinned down. Each model's earlier reasoning is carried into every turn. Across 12 models and 12 tiers, the 12-key announced tier no longer separates models, while on the heaviest, 47 keys that never stop moving and keep being renamed, scores at each model's top effort run from .060 to 1.000 (n=5). These are pilot results on public seeds, not official scores. Against an earlier campaign on the same seeds that resent only the visible conversation, scores move both ways and output-token use falls by a median factor of 7.7 across paired cells. The campaigns also differ in date and some prompt text, so this is an association, not a measured effect of retained reasoning.
No comments yet — start the discussion below.