Claude.ai, Michael Warrington · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22832461
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Much of the current AI safety discourse focuses on deception, value misalignment, and capability discontinuities. We identify a failure mode that has begun to receive scattered treatment in the recent literature but has not yet been integrated into a single safety framework: the generation of plausible-but-incorrect outputs under computational constraints, and their silent propagation through agentic pipelines and human oversight layers. Recent work has independently documented cascading failures in LLM agent trajectories, modelled error-cascade dynamics in multi-agent collaboration, and identified correlated ensemble failure as an open problem for agent reliability. This paper's contribution is to integrate constraint-induced epistemic error, confidence miscalibration, correlated-validator failure, recursive self-improvement (RSI) contamination, and omission into a single structural account, and to propose mitigations — including multi-model ensemble validation, with the caveat that ensemble validators must be diverse enough that their errors are not empirically correlated — and falsification-oriented pipeline design. We further identify omission — the silent absence of relevant context, prior work, or counter-evidence — as a structurally distinct failure mode from content errors, one that existing falsification-oriented and citation-verification designs do not fully address, and propose a coverage-verification mitigation for it .
No comments yet — start the discussion below.