Nathan Ryan Young · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23094709
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Linear readout instruments -- logit lenses, tuned lenses, Jacobian lenses -- decode a language model's residual stream through its unembedding. This paper shows that the component of mid-layer computation these instruments structurally discard is (1) the large majority of the write by energy, (2) load-bearing in aggregate, (3) thin in corpus-prediction content but heavy with algorithmic working state, and (4) READABLE: a linear probe trained on the discarded component alone decodes the model's in-context algorithm at 83.1 percent top-1 (chance 1.6 percent), with leakage and forward-state controls at floor. We call this component the model's computational dark matter, and the analogy is quantitative, not decorative: like its astronomical namesake it is diffuse rather than localized (a halo, not a vault -- the hidden-critical-layer hypothesis is falsified twice), it is weighed by its effects rather than seen (joint ablation), and its darkness is a property of the instrument, not the substance. On a modern 7B model the phenomenon is stronger: readout-visibility of mid-stack writes is approximately zero through two-thirds of the depth -- and there the contrast completes itself: the probe reads the fully-dark component at 12x chance while the lens-aligned component reads at chance (1.8 vs 1.6 percent). The boundary is reported honestly: a plan-probe with context baselines finds working state and the forming next token in the halo, but no readable far-horizon plan in a small non-reasoning model -- the pre-registered null that gives the program's next experiment (reading deliberation in a reasoning model) its meaning; the first two deliberation verdicts on a distilled reasoning model are reported ahead of their cross-model verification cycle and scoped as such. All experiments were pre-registered with sealed gates before code was written; kills and self-caught design flaws are reported at full weight. Every result reproduces on a single consumer GPU in minutes from released scripts. Author contribution and use of AI: the research program and claims are the author's; experiments, derivations, and drafting were done in collaboration with Claude, an AI system by Anthropic, under the author's direction, who verified the results and is responsible for the work. See the corresponding section in the PDF.
No comments yet — start the discussion below.