Xu Zhou · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23118264
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Query perturbations can change retrieved evidence without changing the intended question. We report a frozen Natural Questions (NQ) study combining manually screened query pairs, BM25s top-𝑘 set/rank diagnostics, a lexical evidence-availability proxy, downstream EM/F1, and paired comparisons of four cached query/context inputs: 𝑄𝑐 + 𝐷𝑐 , 𝑄𝑝 + 𝐷𝑝 , 𝑄𝑝 + 𝐷𝑐 , and 𝑄𝑐 +𝐷𝑝 . The balanced primary cohort contains 100 pairs (25 per perturbation family), selected by frozen rules from 185 pairs labeled ACCEPT in manual semantic-preservation review. Its top-ranked passage changed in 50%. The answer-alias-in-passage@5 lexical proxy was 0.47 for clean versus 0.41 for perturbed queries (perturbed-minus-clean difference −0.06, 95% CI [−0.13, 0.01]). The same 100 pairs underwent 400 DeepSeek generation requests. Exact match was 0.38 under 𝑄𝑐 + 𝐷𝑐 and 0.36 under 𝑄𝑝 + 𝐷𝑝 ; the prespecified equal-family-weighted point difference was 0.02, while the frozen pooled base-query 95% bootstrap resampling interval was [−0.06, 0.10]. All overall query-surface, evidence, and interaction contrast intervals included zero. Separately, a 500-base/2,000-pair automatic-candidate retrieval diagnostic, whose inclusion did not require manual semantic-review acceptance, showed a top-1 passage change in 52.7% of pairs, mean Jaccard@5 of 0.428, and a lexical proxy of 0.474 versus 0.3685. The two analyses have different membership and validation status. Retrieval-list instability is observed; the final-100 lexical proxy and downstream answer-score differences are imprecisely estimated and do not resolve an overall decrease.
No comments yet — start the discussion below.