Keerthi Sagar Chegondi · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22859445
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Offline evaluation can overstate machine-learning performance when features, transformations, records, or evaluation design expose information that will not exist at the production scoring moment. LeakageBench is an applied benchmark and release-control framework for prediction-time integrity. The study defines a fixed pre-decision scoring boundary, classifies model inputs by their deployment-time availability, compares chronological and random evaluation designs, tests multiple leakage mechanisms, and independently reconciles released metric effects before public claims are approved. The goal is not to introduce a new predictive algorithm; it is to separate model performance from information leakage and to make the evidence behind a pass, warning, or block decision auditable. The resulting case study demonstrates how deployment-valid evaluation, feature contracts, reproducible evidence, and claim governance can be joined into one machine-learning release workflow. **Keywords:** prediction-time integrity, data leakage, machine-learning evaluation, temporal validation, model governance, deployment validation, reproducibility, feature availability ---
No comments yet — start the discussion below.