Nicholas Templeman · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22985467
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Evaluation results for large language models are usually published as totals. We ask whether those results are a property of the model, the prompts and the grader, or also of the machine that ran them. We re-ran 154 published evaluation results (cards), covering 11 open-weight models on 14 evaluation axes and 9,757 items, on a second runtime (a Kaggle Tesla T4 instead of the RunPod RTX 3090 that produced them), with the same model manifest digest, prompts, item banks, instrument and greedy decoding. Raw outputs were byte-identical on 84.4% of items; grades agreed on 97.3%. The 260 grade flips (2.66%, 95% CI 2.36–3.00%) went in both directions (122 up, 138 down) and mostly cancel in totals: 62 cards reproduced their total exactly, but on 11 of those an item's grade had changed. Only 54 cards were grade-identical item by item. A further 83 items swapped between parse error and wrong answer, changing the reported accuracy's denominator. Re-running the T4 job on the same T4 was not byte-deterministic for 20 of 150 cards. We argue that evaluation results should be admitted for publication only on item-level agreement across independent runtimes, and that the runtime must be reported with every result. Per-item data, runtime declarations, signed Merkle roots and an offline verifier: https://huggingface.co/datasets/csoai/cross-hardware-reproducibility (CC-BY-4.0). Measurement, not certification: no model, company, GPU vendor or cloud provider is scored.
No comments yet — start the discussion below.