Adam Allcock · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23048151
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Frontier language models can be highly accurate on a task yet fail to perform it reliably enough to run unattended. On long chains of simple, fully specified clerical computation, such as billing, payroll and ledger replay, the binding constraint is not capability but consistency, and ClerkBench measures both as workload lengths. The capability horizon is the length at which a single attempt's probability of producing a byte-exact artifact falls to 50%; the headline consistency horizon is the length at which succeeding on all six of six attempts falls to 50%. Models are tested without tools, so the measurement is the model's own execution reliability. Each instance is generated from a seed and graded against a simulator; the scored set is frozen and released in full. Across fourteen configurations from four frontier laboratories, horizons run from below the smallest rung (50 events) to beyond the largest (400). GPT-6 Astra and Claude Opus 5.5 outrun the ladder on at least one family, so their scores are lower bounds (≥400 and ≥317). Among located horizons GPT-5.5 leads at 147 events, ahead of DeepSeek V4.1 Flash at 138 for about a tenth of the price per attempt, but the leading scores' bootstrap intervals overlap. Matched controls rule out a pure copy or serialization bottleneck: the horizon measures computation, though the score is operational rather than mechanistic. Repeating each task twenty times separates two kinds of failure: random slips, whose consistency cost follows from single-run accuracy, and model–instance traps, where a model fails one instance 16 times in 20 while near-identical siblings pass 17 and 19 of 20. Traps are invisible to single-run scores and did not transfer between the models tested, so consistency must be measured, not inferred. Code and data: github.com/adamallcock/clerkbench, archived as doi:10.5281/zenodo.23048121.
No comments yet — start the discussion below.