Nathan Ryan Young · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23094715
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This paper reports that "how hard a language model is thinking" is a measurable physical quantity: the total second-order cross-curvature of the output margin at the model's own binding site. The instrument is a mixed-second-difference kernel in true fp32 (bf16 is blind for second-order work), validated by eps-squared scaling across five independent runs. Findings, each pre-registered: derivation carries a curvature surcharge over recall (1.66x, CI [1.42, 1.92]) that behavioral accuracy cannot see; the surcharge sits at a model-specific BINDING SITE (source-side at the entity token on the 7B, destination-side at the answer slot on the 1.5B/3B/14B -- the rule is "bind where the last ingredient lands", and prompt geometry can MOVE the site); the meter is GRADED and calibrated in hops (in-context chains: Q(k) = 1.00/1.68/2.14/3.14, monotone, CI-solid; two-hop chains and two-hop geography price identically, 1.68 vs 1.66); the schedule is FIXED (derivation never costs layers -- arrival depth is task-invariant within every family tested, and the deadline setting is a family property: 0.89/0.75/0.69 of depth); chain-of-thought CONVERTS work rather than abolishing it (the derivation site goes quiet, the answer slot pays retrieval, total discount ~23 percent); and error classes are REAL and differentially treatable -- transient-gold errors (the right answer alive mid-stack, lost at shipping) are rescued at up to 79 percent by amplifying the model's own gold direction in the shipping band (the AUTOGRAFT: the cure is the model's own tissue), with random-direction controls at zero and a per-model therapeutic window. The meter's honest limits are part of the result: it is a DEMAND gauge, not a truth gauge (per-item work does not predict correctness, AUC 0.44); Mistral shows no kernel signature at any probed query-side site (a shaped null; the eager-binder discriminator is reported); cross-family metering requires per-strategy site calibration. Battery status (2026-07-11): the launch-wave arms all landed and are folded into the draft -- the statement-position discriminator (cold: Mistral shows no kernel-visible site anywhere probed), the conservation arm (~23 percent discount, the rest converted), the honest-meter null, and the cross-model rescue window; still owed before upload: a third metered family, a third task family, seed batteries on the site sweeps, and this paper's own adversarial audit rounds. All experiments pre-registered with sealed gates committed before code; misses of the sealed predictions are scored in the text. Author contribution and use of AI: the research program and claims are the author's; experiments, derivations, and drafting were done in collaboration with Claude, an AI system by Anthropic, under the author's direction, who verified the results and is responsible for the work. See the corresponding section in the PDF.
No comments yet — start the discussion below.