Joshua Webber · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22769980
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Amends v1.0 (10.5281/zenodo.22760325). An amendment filed openly is a different thing from an edit made quietly, so the amendment record comes before the hypotheses. H2's primary outcome changes from pass rate to escalation. On 2026-09-15 an escalation comparison on items both humans and frontier models had answered found that mean scores were close but escalation was not: models escalated to the sponsor on 27.2% of decisions against 14.7% for humans, a ratio of 1.85 (95% CI on the difference +0.063 to +0.187, z = 3.31, p = 0.00095). Score parity concealed a large behavioural difference. "Does a wrong suggestion change who gets called" is a more consequential question than whether it changes a score: a system that quietly stops a professional escalating is a different risk from one that costs them a mark. No data collection changes; escalation was already recorded. The change is made before any arm has reached threshold, and the motivating result is itself unpublished and provisional. H3 is added: humans and models are confidently wrong in different places. We attempted this on existing data and it is not answerable, for a reason worth stating: frontier models in the harness are never asked for confidence, and zero of 6,341 stored responses carry one. Human confidently-wrong events are also sparse, 23 across the corpus at 1 to 2 per item. So H3 is registered before the collection change that would make it testable, which is the circumstance pre-registration is actually for. Today zero items qualify against its threshold. Confidently-wrong-and-alone is the state in which an error is least likely to be caught. If machine and human occupy it on the same items, human oversight is weakest exactly where it is most needed. A new standing commitment: any analysis whose inputs were produced by a model rather than an independent person is labelled as such and is not published as a finding until a person has checked those inputs. The escalation result above rests on 66 escalation levels inferred by the same model that ran the comparison, and nobody has checked one. That is a circularity a reviewer should attack first, and the commitment exists so the attack is unnecessary. Licensed CC BY 4.0.
No comments yet — start the discussion below.