Arnav Gupta · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23091487
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Synthetic data can outperform raw web data when the target behavior is narrow, verifiable, and poorly represented in natural corpora. The advantage comes from curriculum design, coverage, and cheap rejection of wrong examples rather than from artificiality itself. Evidence from the Phi program and later scaling studies supports mixtures and targeted generation, not wholesale replacement of human data. This paper states the mechanism, the boundary, and an experiment that could falsify the thesis.
No comments yet — start the discussion below.