László Fazakas · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22758471
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Four mechanisms are used to make machine-assisted decisions trustworthy: a formally verified derivation, a panel of independent models, an evaluation benchmark, and a regulatory oversight duty. Each is better than its absence, and each carries a presupposition it does not itself check — that the type system was the right carving of the domain, that panel members do not share a failure mode, that the reference was authored independently of what it grades, and that the supervisor retains the capacities supervision requires. This paper argues that these are one limitation rather than four: the checking mechanism inherits a presupposition it does not check, so two materially different states of the world produce one observable, and the instrument built to distinguish them is the instrument that cannot. Six instances are set out with their evidence graded separately, including one in which the author's own published headline result did not survive its control. The consequence that makes the structure worth naming is that improvements in reasoning capability widen these gaps rather than closing them: a stronger derivation attracts less scrutiny, a more fluent answer conceals a substitution more completely, better agreement is worse evidence when the prior is shared, and a training loop indexed on measured gaps improves only the failures that already have an instrument. Six refutation conditions, three remedies that do not depend on the capacity in question, and five open problems are stated. Position paper, draft v0.2. Not peer reviewed. Circulated for refutation. Competing interest declared in the author note.
No comments yet — start the discussion below.