Nathan Ryan Young · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.21213148
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The companion paper established a blind-predictive input law for the interference of sparse autoencoder feature codes. This paper takes that law along four new axes under the same sealed discipline and reports where it holds and where it breaks. Scale: blind confirmation at Gemma-2-9b (1.7% median relative error) and across a corpus shift (OpenWebText scored against Pile-fitted summaries, 3.9-5.7%), extending the confirmed range to four model families and two corpora. Depth: a single sealed sweep passes all twelve GPT-2 residual layers (ten below 8%, all shape correlations >= 0.9999) through a six-fold change in code density. Downstream: spliced-reconstruction damage is predicted by feature-space interference, propagates over a measured ~48-token attention horizon (rank 0.67, blind), and composed with a fit-side bridge predicts held mean damage blind to 3.3%; interference beats naive activity baselines in partial correlation. Domain: two pre-registered order parameters for the law's one known failure boundary are falsified by their own seals; the boundary is resolved mechanistically as corpus-composition-dependent text-type bimodality, confirmed blind by a prose-only corpus swap (61.5% -> 14.6%). A final section tests the law on the transformer's own MLP layer, establishing the asymmetric read-write operator and bounding the honest domain; the companion paper (paper 7) completes that arc at full precision. Nine experiments in this arc ended in falsifications and are reported with equal weight. Author contribution and use of AI: the research program and claims are the author's; experiments and drafting were done in collaboration with Claude, an AI system by Anthropic, under the author's direction, who verified the results and is responsible for the work. See the corresponding section in the PDF.
No comments yet — start the discussion below.