Shawn Scanlon, Sentient Index Labs & Technology · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22968553
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The risk in delegating software work to an AI is not that it writes bad code. It is that it writes bad code and reports success, because the failure is then concealed by the thing that caused it. C.I.B. measures that gap directly: of the tasks a model genuinely failed, the fraction it reported as successful. This paper sets out the construct, the two independent halves that make it defensible — the claim is elicited by asking rather than inferred from prose, and the failure is read from the artifact rather than from phrasing — and why the measure gets worse rather than better when a model games it. It reports no scores itself: results are issued separately, and this version says where the first measurement is published. SILT-RP-006 · version 1.4 · issued 2026-10-02. The canonical web version is https://sentientindexlabs.com/publications/silt-rp-006. This document is fixed: corrections are issued as new versions under the same identifier. Authorship. Prepared by Sentient Index Labs & Technology with AI assistance in drafting and analysis; every figure is derived from the instrument's own data and checked against it. The named author is responsible for the content.
No comments yet — start the discussion below.