Yipeng Ma · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23000327
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
A negative result, reported in full. Adding five domain-derived features plus a dispersion index to a Double Machine Learning (DML) pipeline reduced the bias of the estimated average treatment effect (ATE) by 73.75% and collapsed the estimator's standard deviation by 4.00× — a result that would ordinarily be reported as a methodological advance. This study shows those gains were entirely attributable to target leakage: one derived feature was a deterministic function of the untreated potential outcome Y(0), which is unobservable by construction. Using a controlled ablation that holds the feature set fixed and toggles only the leakage channel, repeated over 30 independent random seeds, the same features yield a bias change of −0.65% (paired t = −0.137, p = 0.892; 16/30 wins) once leakage is removed. The apparent improvement disappears completely. The study further characterises the empirical signature of leakage. Sweeping leakage intensity α ∈ [0,1] over 11 levels, bias decreases monotonically while estimator dispersion collapses by 4.00× and the reported OLS standard error collapses by 4.65×. It is argued that a simultaneous large reduction in both bias and variance — particularly in reported standard errors — should be treated as a warning signal rather than evidence of success. A second, independent class of error in the same pipeline is documented: misdefinition of the ground-truth ATE, where using the control-group mean introduces a 3.53% error and treating the naive difference-in-means as theoretical truth inflates the target by 81.16%. Findings are distilled into a diagnostic checklist for practitioners building feature sets for causal inference. This deposit contains: the manuscript (PDF and LaTeX source), the full reproduction code (src/), cached results (results/), a self-contained executed demonstration notebook (notebooks/), and figure sources.
No comments yet — start the discussion below.