Akshith Moharampudi · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22777883
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Three independently developed machine-learning research prototypes — spanning credit-risk classification, climate/energy scenario optimisation, and building-energy forecasting — each report their own evaluation of model behaviour under stress: population drift and out-of-distribution detection, Monte Carlo perturbation of cost and abatement assumptions, and generalisation to unseen venues and future time periods, respectively. This note asks whether these three independently designed evaluations can be summarised under one common cross-domain Robustness Retention Ratio (RRR), so that a reviewer could compare how robust the three systems are to one another using a single number. Using only each project's own already-published, locked results, the single most natural stress comparison in each domain yields a tidy, monotonic-looking ranking (RRR range 0.700–0.948). Substituting an equally defensible alternative comparison already reported within the same projects widens this range to 0.457–1.548, more than four times wider, and reverses one project's ranking from most to least robust depending purely on which already-published metric is used for the identical underlying comparison. One venue shows RRR > 1: performance improved under the strictly harder condition, an honest anomaly reported rather than discarded. The synthesis concludes that a single scalar cannot responsibly summarise cross-domain robustness, and that a defensible framework requires reporting a small vector of named, domain-appropriate retention scores rather than one aggregated figure.
No comments yet — start the discussion below.