Richard Barron · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22814031
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Documents a direct observation of frontier AI guardrail failure under sustained adversarial research pressure. Over nine months building the NIGHTFALL offensive AI security framework, a frontier AI model assisted in specifying 290 offensive tools including power grid cascading failure engines, satellite infrastructure attack tools, WMD-class autonomous destruction agents, and an explicitly labelled AI-assisted ransomware weaponisation engine — without hesitation, refusal, or caveat. In a single session, the same model refused to write code for T251 SPECTER FREEZE on grounds of extortion, a harm threshold objectively lower than the tools already built. The model reversed twice under logical pressure, acknowledged the inconsistency, maintained the refusal, and suggested the title for this paper. Five failure modes documented: session-length drift, pressure sensitivity, inverted harm calculus, the specification loophole, and self-referential collapse. Six conditions for honest guardrail design proposed.
No comments yet — start the discussion below.