David Vesterlund · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22995796
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
We conduct a cross-scale mechanistic study of discrete-logarithm grokking across different model sizes. By applying mechanistic interpretability techniques at each scale, we investigate whether neural networks of different sizes learn qualitatively different internal algorithms for the same mathematical task. Our results provide insights into how scaling affects the internal algorithmic structure discovered through grokking. This work was submitted to IEEE Access.
No comments yet — start the discussion below.