Michael Warrington, Claude.ai · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.21862079
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Debate has been proposed as a scalable oversight mechanism for evaluating AI outputs that exceed the checking capacity of a human or weaker judge, on the theoretical grounds that refuting a false claim is easier than generating a convincing one. Empirical work on this claim has been conducted almost entirely on tasks where ground truth exists but is withheld from the judge (and often from the debaters). This leaves open a distinct and more consequential question: does debate produce reliable verdicts when no ground truth exists for anyone at judgment time — the actual condition under which novel scientific hypotheses are generated? This paper proposes a protocol for testing that question directly, rather than assuming an answer. New mathematical or structural results are applied to existing datasets to generate candidate hypotheses; a diverse population of AI systems is first asked to answer independently and blind, with the resulting agreement pattern used to flag which hypotheses need scrutiny and which may be shared training artefacts; contested hypotheses then proceed to adversarial debate under asymmetric roles; verdicts are recorded and sealed before any external validation is available. Ground truth is allowed to arrive later, through independent replication, downstream falsification, or convergence with subsequent findings, and the sealed verdicts are then scored against it. The contribution is not a new debate mechanism but an empirical test of whether an existing one is trustworthy precisely where it is currently unverifiable.
No comments yet — start the discussion below.