Pranay M. Mahendrakar · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22840564
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. A retrieval-augmented system that answers only when its retrieved context suffices needs a predicate that says when it does. Three incompatible predicates are in circulation under the word "sufficient", and results obtained under different ones are routinely compared as though they measured one quantity. The first is answer-free: an instance has sufficient context if some plausible answer to the question exists given the information in the context, a definition whose authors state explicitly that it does not require a ground truth answer and that it permits the context to contain an incorrect answer. The second is answer-conditional: this evidence supports this candidate claim, the SUPPORTED, REFUTED and NOT ENOUGH INFO structure inherited from fact verification, whose framing document says in as many words that it avoids absolute judgements of factuality. The third is gold-relative: the context supports the gold answer, which is what unanswerable-contrast dataset construction has meant since 2018. This paper argues that only the third predicate entails correctness, that the third is the only one of the three that cannot be computed at inference time because it needs the answer the gate exists to decide whether to produce, and that the decoupling the field reports as an empirical surprise is in one direction a consequence of the definition that makes the gate deployable at all. The published measurements are consistent with that reading and are not usually read that way. On a curated set carrying human-annotated sufficiency labels rather than machine ones, three frontier models hallucinate on 3.2 to 14.3 percent of sufficient-context instances and a 27 billion parameter open model on 25.4 percent, while under insufficient context on that same set they are still correct 7.7 to 23.1 percent of the time; across the larger autorater-labelled analysis the insufficient-context correctness rate runs from 35 to 62 percent, and the study's own qualitative table attributes it to eight causes of which parametric knowledge is one and outright guessing on yes/no and limited-choice questions is another. Both sides of the relation are measured by language models, and each instrument moves the number by double digits: the sufficiency autorater is validated at 0.930 accuracy on 115 instances drawn from four datasets, two of which are not the datasets it is then applied to, and switching the correctness metric from lexical containment to a model judge moves one insufficient-context figure from 46.1 to 59.5 percent and one sufficient-context figure from 48.9 to 74.0. The behaviour the gate is meant to correct does not track sufficiency either: abstention falls when retrieval is added, from 84.1 to 52 percent for one model and from 100 to 18.6 percent for another, and a 2026 controlled study of three small open models finds abstention keyed on whether context is present rather than on whether it supports anything, with answer rates under should-abstain conditions rising from 3.0 percent when context is missing to 82.7 percent when it is misleading. The premise underneath is itself disputed: a 2024 result that adding random documents improves accuracy by up to 35 percent was reproduced in 2026 and found not to survive changes to prompting and decoding. Sufficiency-shaped gates do work where they are built and measured, and the two strongest 2026 systems are read here at their best rather than their worst. What is missing is the joint table: no located work reports all three sufficiency predicates on the same instances, so the rate at which they disagree is unmeasured, and the decision each licenses is therefore unlicensed. The statistical machinery for that table was published in 2026 and carries a single retrieval-success variable; splitting it three ways is the experiment this paper asks for. The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the DataCite or Crossref record for its DOI before inclusion, and every quantitative claim was read back against the cited source's own table or text before it was written down. No experiment was run and no number in this paper was measured by its author; every number is quoted from the paper credited with it, except for one confidence interval computed from a published accuracy and sample size, which is labelled as a calculation at the point where it appears, and for differences between two published figures, which are stated as differences. Section 2 states the search procedure and its limits so that the coverage claims in Sections 11 and 12 can be checked and, if wrong, corrected. The author is responsible for the final text and for all claims made in it.
No comments yet — start the discussion below.