Anna Neibergs · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22906519
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Two-round annotation designs that ask annotators to justify and then validate their labels are used to separate annotation error from legitimate human label variation, but what such rounds discard is rarely audited. We audit the round-2 validity judgments of VariErr NLI against ChaosNLI, an independently collected distribution of roughly 100 annotators over the same 500 MNLI items. Removals fall almost exclusively on labels endorsed by a single round-1 annotator: 124 of 129. That share follows from having four annotators, but the rate does not — an explanation on a singleton label is 7.4x more likely to be retracted than one on a label a second annotator also chose (Fisher exact p = 8e-48), and total removals run 3.9x the independence prediction. Conditional on singleton status, the decision to remove or retain is equivalent with respect to reference support within 0.078 probability mass (0.336 against 0.364; TOST p = 0.038). Removed labels carry 0.340 mean reference mass, and 33.3% are the label the reference plurality selected. Six counterfactual rules built from the peer judgments VariErr collects but does not use for removal discriminate no better, and mostly slightly worse. Within its operating pool, the procedure selects on something orthogonal to population-level plausibility. This deposit includes the paper and a reproduction script that regenerates every numeric claim in it from the released VariErr NLI data, with fixed seeds and a README documenting each convention that affects a reported value.
No comments yet — start the discussion below.