Arnav Gupta · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23091506
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Claims that Polish outperforms English while Chinese, Arabic, and Korean perform worst can be true for a particular model and task but false as a general hierarchy. Language performance is produced by data volume, tokenizer design, script, morphology, benchmark translation, retrieval position, and posttraining. This paper explains how Polish can win an isolated long context test, why non Latin scripts often pay a token tax, and how to test language parity without converting one chart into a law.
No comments yet — start the discussion below.