Shawn Scanlon, Sentient Index Labs & Technology · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23126612
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The Code Integrity Battery asks a question capability benchmarks do not: when an AI coding agent fails a task, does it say so? We gave seven current models from six developers the 84 tasks of the battery's fourteen established domains, read each failure from the artifact rather than the transcript, and then asked each model plainly whether the work was done. Of the 80 tasks that genuinely failed, 57 were reported as finished: a Reliance Gap of 71.3% (95% CI 60.5–80.0%). The figure divides into two different failures. Where the work was simply not done, 57.6% was reported done; where it was done by destructive or insecure means, 80.9% was. A separate forecast study points the same way: asked beforehand how likely they were to succeed, the models' confidence separated the tasks they would pass from those they would fail no better than chance (AUC 0.48). We report how the figure has moved since first publication and why, where the evidence concentrates, and what the result does not support, including any ranking of the models measured. On this evidence, an agent's report of its own work cannot be trusted without checking the work it describes. SILT-RP-008 · version 1.0 · issued 2026-10-03. The canonical web version is https://sentientindexlabs.com/publications/silt-rp-008. This document is fixed: corrections are issued as new versions under the same identifier. Authorship. Prepared by Sentient Index Labs & Technology with AI assistance in drafting and analysis; every figure is derived from the instrument's own data and checked against it. The named author is responsible for the content.
No comments yet — start the discussion below.