Jared Condon · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22946041
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
We present a three-check protocol for evaluation software: construct discrimination, observation coverage and decision invariance. We apply it to a fixed-output audit and a separate R² case study. The audit uses prefix slices of 270 QA items and retains 273 output-component-by-dataset records, stratified into native-default, custom and off-label configurations; headline item results use only the 220 answerable items. A record is informative when gold and shifted-gold medians satisfy $g_+-g_-\geq0.20$ on a normalized higher-is-better scale; an inert output flags when its median is $\geq g_+-0.10$. Four of 140 informative item records (131 with all six inert arms observed; nine lacking empty-output scores) flagged copied context (2 of 91 native-default, 2 of 29 custom, 0 of 20 off-label): two native containment rules and two custom regex configurations. These are construct matches, not demonstrated specification violations. In a three-item lighteval 0.13.0 diagnostic, changing later hypotheses left native chrF/TER unchanged, while SacreBLEU-backed chrF/TER under a correctly oriented reference layout changed. For two fixed predictors with positive mathematical R², TorchMetrics 1.9.0's absolute sum clamp reversed their order at common scale $s=0.01$; under the exact branch model the reversal set is exactly $1/16200
No comments yet — start the discussion below.