Ravil Akhtyamov · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23061304
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
We introduce KHWAREZM-100, an evaluation benchmark for large language models (LLMs) built from primary sources of the medieval Khwarezm mathematical school (al-Khwarizmi, al-Biruni, al-Kashi), and use it to argue for a previously under-recognized class of LLM evaluations we term access-bounded benchmarks: tasks on which the frontier gap is bounded below by institutional and linguistic access to training data, not by model scale or compute. Our central contribution is a formal one. We model the frontier laboratory's corpus-acquisition decision as a budget-constrained allocation problem and prove, under a support-conditioned scaling law and a non-substitutability assumption, that the resulting benchmark gap admits a lower bound delta > 0 that is uniform in compute (Theorem 1); no finite compute multiplier closes it (Corollary 1). The proof separates two mechanisms that the informal literature conflates - hard bounds, where the corpus does not exist in machine-readable form, and soft bounds, where it exists but its acquisition is never optimal for the frontier - and this separation resolves the apparent paradox of releasing an access-bounded resource openly (Section 2.4). Three further contributions follow. (i) We instantiate the category with KHWAREZM-100, drawn from the Kitab al-jabr wa-l-muqabala and related corpora, scoring each item jointly on solution correctness and methodological fidelity - adherence to the historical six-case taxonomy and completing-the-square geometry - a dimension absent from standard mathematical benchmarks and made concrete by two worked derivations (Appendix E). (ii) We specify the estimation protocol quantitatively: a closed form for the corpus-sufficiency point alpha*, and a power analysis establishing that the fame-fidelity correlation claim requires n >= 33 source-locked items, which our pilot (n = 15) does not meet (Section 4). (iii) We propose a tri-lingual sentence-aligned (Arabic-Russian-Uzbek) TEI-XML corpus of these sources for targeted pretraining. As preliminary empirical content we report partial construction and a single-model correctness probe with a blind, no-LLM-judge harness; the multi-model, human-scored fidelity study is in progress.
No comments yet — start the discussion below.