Thomas Edrington · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.21753454
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
We present a systematic methodology for falsifying mechanistic interpretability findings before publication, developed across six months and three research programs on transformer geometry, KV-cache phenomenology, and workspace selectivity. We document 19 cases that would have survived standard peer review -- killing 11 outright, correcting 3 whose metrics were flawed but whose underlying findings survived with proper baselines, partially killing 2, finding 1 inconclusive due to measurement-methodology disagreement, negatively replicating 1 external prediction, and flagging 1 for limited generalizability. The paper is the methodology as much as the findings: pre-registered confound checks, multi-baseline controls, and adversarial replication that catches what reviewer intuition misses. The lesson is uncomfortable: the bar for a 'real' mechanistic finding is much higher than the bar for publication.
No comments yet — start the discussion below.