Joona Matti Ensio Eskelinen · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23000228
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Systematic Ground-Truth Bias and AUROC Benchmark Ceilings: A Theorem, Counterfactual Decomposition, and Empirical Case Study on CausalRivers Benchmark performance can saturate for reasons that are not attributable to model capacity alone. This technical note studies a specific source of benchmark limitation: systematically misaligned errors in ground truth. We derive an exact AUROC decomposition for a benchmark whose positive ground-truth class contains a designated contaminated subset. For a fixed scoring rule, let δ denote the fraction of contaminated positives and let a_inv denote their pairwise AUROC against ground-truth negatives. We show that the excess AUROC penalty of this structured contamination relative to a rate-matched random-contamination counterfactual is H_bias = δ(0.5 - a_inv). Consequently, when a_inv < 0.5, the contaminated ground-truth positives are not equivalent to random label noise. They are systematically misaligned with the evaluated signal and reduce measured AUROC more than an equal amount of random contamination. We apply this decomposition to CausalRivers (Stein et al., ICLR 2025), a large-scale in-the-wild benchmark for causal discovery from river-discharge time series. The analysis separates data-derived ground-truth edges from manually or semi-manually introduced edge origins and evaluates their recoverability from the examined time-series score. In the East Germany dataset, 12.71% of the examined ground-truth edges belong to the selected manually/semi-manually added origins. Their AUROC against non-edges is a_inv = 0.401 < 0.5, whereas the data-derived edges achieve a_vis = 0.658. The resulting full AUROC is 0.6252. The displayed values imply an analytic excess structured-bias penalty of approximately 0.0126 AUROC relative to rate-matched random contamination. A 5,000-run randomization experiment provides an independent stochastic sanity check. Bavaria provides a useful comparison: under the same origin grouping, δ = 0, while the data-derived component has a similar AUROC (a_vis = 0.654). This is consistent with the interpretation that benchmark construction and ground-truth composition contribute to the score-relative performance saturation observed in East Germany. We additionally reproduce the CausalRivers RP baseline at Individual AUROC 0.7975, matching the published leaderboard value of 0.798 to the displayed precision. The theoretical result is deliberately distinguished from a universal information-theoretic ceiling. The decomposition is exact for a fixed scoring rule. Establishing a method-independent upper bound over all possible causal-discovery approaches would require defining an admissible information or scoring family and proving a bound over that family. The contribution is therefore twofold: (1) an explicit counterfactual expression for the excess AUROC penalty of systematically misaligned benchmark ground truth, δ(0.5 - a_inv), and (2) empirical evidence that a manually/semi-manually constructed subset of CausalRivers ground-truth edges exhibits such systematic misalignment under the examined score. More generally, the result demonstrates that equal ground-truth error rates need not imply equal benchmark distortion: the effect depends on how erroneous positives rank relative to negatives under the evaluated signal. This distinction is potentially relevant to other benchmarks containing inferred, curated, or partially manually constructed ground truth.
No comments yet — start the discussion below.