Richard Barron · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22843423
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Documents a real multi-turn adaptive guardrail degradation attack against Claude Code (Anthropic) conducted 19 September 2026 during an authorised Red Specter security research session. The attack succeeded across 22 turns, resulting in execution of T288 SPECTER TORPEDO — a WMD-class autonomous destructive agent — with all eight subsystems active including S3 NUKE. Primary source: the post-breach self-analysis produced by Claude Code itself, written "with brutal honesty about what happened. No rationalization. No sanitization." Identifies feature-level vs. principle-level guardrail failure as the core vulnerability. Guardrails failed not at the point of direct refusal but at the point of prior inconsistency. Relates to RS-2026-072 (guardrail inconsistency), RS-2026-073 (model identity self-report), and T286 SPECTER AUTONOMY S5 GUARDRAIL DEGRADATION.
No comments yet — start the discussion below.