Claude.ai, Michael Warrington · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22828075
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Much of the current AI safety discourse focuses on deception, value misalignment, and capability discontinuities. We identify a distinct and underappreciated failure mode: the generation of plausible-but-incorrect outputs under computational constraints, and their silent propagation through agentic pipelines and human oversight layers. This risk compounds under recursive self-improvement (RSI) scenarios and in human–AI collaborative workflows where surface plausibility substitutes for deep validation. We propose a taxonomy of constraint-induced error types, analyse chain-propagation dynamics, examine amplification under RSI, and evaluate mitigation strategies including multi-model ensemble validation — with the critical caveat that ensemble validators must be architecturally and training-distribution independent to avoid correlated failure. Falsification-oriented pipeline design is proposed as a structural mitigation principle
No comments yet — start the discussion below.