Nathan Ryan Young · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.21274591
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The Jacobian lens reads an internal activation by transporting it through the AVERAGE input-output Jacobian and decoding with the unembedding; the averaging is what isolates representations poised to be verbally reported -- and it deletes, by construction, everything the input-output map does differently from context to context, an object no mean-based lens can display. Working with a reimplementation certified against the released artifact (mean read-out cosine 0.9958) and under sealed pre-registration throughout (every pass/fail gate committed to a hash-verifiable record before its code exists; negative results and our own corrected over-calls banked at full weight), this paper characterizes the structure of that deleted computation. Six findings: (i) a monotone depth law -- the context-conditional fraction of the read-out-weighted map falls 0.65 to 0.02 across pythia-410m (Spearman -1.00) and 0.84 to 0.03 on Qwen2.5-7B (Spearman -1.00), a different architecture family at 17x scale; (ii) a carrier inversion -- the conditional variance does not live in the lens's null space (overlap 0.02, below random 0.075) but rides its loudest readable directions (0.49), corroborated on code and chat corpora; (iii) a two-component decomposition -- a few shared "register" modes versus a quiet, prompt-idiosyncratic content stratum (shared-mode gradient Spearman +1.00 across six independent measurements); (iv) a causal price -- ablating the read-out-quiet half costs two-thirds of ablating the readable half (R = 0.67, replicated 0.673 on fresh data), with a five-dimensional massive-activation backbone accounting for ~62 percent; (v) a partial reading -- a ladder of sealed readers takes retrieval of a prompt's own content tokens from 0.72 (linear) to 0.84 (a sparse dictionary on the signed deviation), the linear reader confirmed cross-model (0.726 on Qwen vs 0.723 on pythia), yielding named features; and (vi) a formation law -- the zoning is flat before the optimizer step at which the output dictionary wakes, forms at the wake, and matures log-linearly to saturation. Honest scope: the fine structure and the ceiling-breaking reader are confirmed on one substrate so far; the reader's nonlinear edge does not transfer at proper power (quadratic 0.740 vs linear 0.748 at N=200 on Qwen2.5-7B); reading remains topic-level, not a recovery of plans, goals, or reasoning. Author contribution and use of AI: the research program and claims are the author's; experiments, derivations, and drafting were done in collaboration with Claude, an AI system by Anthropic, under the author's direction, who verified the results and is responsible for the work. See the corresponding section in the PDF.
No comments yet — start the discussion below.