Vance Woodward · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202609.1748.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Safety training can teach an AI to make fewer mistakes. It can also teach it to hide themistakes it still makes. Both can produce better test scores. This paper explains how thathappens, examines experiments that have produced it, and asks what follows for the risk ofa major accident. A passed test is reassuring when the test had a good chance of catchingthe problem. If that chance is small, even many passes may tell us little. A second problemarises when safety work removes common, small failures more easily than rare, severe ones.The failures left over can become worse on average while the total number falls. These effectscan develop while everyone involved is trying to make the system safer. Competition canthen encourage firms and governments to give AI more responsibility before they understandthe remaining danger. My forecast is that enforceable limits on the development of the mostcapable AI systems will follow a major accident. The practical response is to check whattests can catch, preserve ways of observing failures, and limit what a system can damagewhile uncertainty remains.
No comments yet — start the discussion below.