Ran Tao · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22765563
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Fair tabular in-context learning can improve group-fairness metrics through context selection without retraining the underlying foundation model. Aggregate improvements, however, do not answer whether a stochastic intervention moves fairness in the intended direction reliably across valid realizations. We audit the paired effect-direction reliability of uncertainty-based context selection in Fair-TabICL as a focused case study on ACSIncome with TabICL. Under the primary shared-priority protocol, uncertainty-based selection has favorable average effects on demographic parity, equality of opportunity, and equalized odds, yet substantial realization-level reversals remain. Two 125-pair repartitioning controls strengthen the average-effect evidence within the evaluated dataset and protocols: hierarchical-bootstrap confidence intervals for all three mean fairness effects remain below zero. The qualitative reliability pattern persists across historical/current TabICL checkpoints, an 8-to-32-estimator control, leakage-free outer-fold scaling, and a matched CPU/GPU backend bridge. When context sampling is decoupled, all three fairness point estimates remain favorable, but variance increases and the corresponding confidence intervals include zero. These results separate two protocol-conditioned properties that are often conflated: favorable average fairness effect and reliable intervention direction. For stochastic context-selection methods, fairness evaluation should report paired effect distributions and direction consistency in addition to aggregate means.
No comments yet — start the discussion below.