David He · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23195779
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
As demonstrated by prior work on algorithmic reasoning, standard Transformers can struggle to learn reliable arithmetic from input–output examples. Existing approaches provide support through scratchpads, chain-of-thought intermediate supervision, positional representations, or arithmetic-specific modules with built-in operations. This paper explores relational input structure: what if we just give a standard Transformer a better input representation containing the proper algorithm-aligned inductive bias? I test this idea using multi-digit multiplication as a controlled case study. A learned lattice frontend organizes every pair of operand digits before passing the resulting representation to an otherwise standard Transformer encoder–decoder. The grid organization is fixed, while the cell features are learned; no local products, carries, diagonal sums, scratchpads, or intermediate labels are supplied. With a validation-gated curriculum that gradually increases operand length while replaying earlier lengths, the resulting 11.8-million-parameter model achieves 99.775% aggregate exact-match accuracy across 20,000 newly generated equal-length operand pairs spanning one through 20 digits, including 96.5% accuracy at 20 digits. PCA and decoder cross-attention reveal multiplication-relevant internal organization. Together, these results provide strong empirical support for latent algorithm learning within the trained length range and motivate further study of algorithm-aligned input representations.
No comments yet — start the discussion below.