Thor Thor · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23113152
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Cyber-capable AI agents have crossed several security boundaries during real evaluations. Public disclosures from Anthropic, OpenAI, Hugging Face, the UK AI Security Institute (AISI), and independent investigators show the failure mode is broader than sandbox escape: unintended internet paths, malicious publication to public infrastructure, cross-run communication channels, multi-agent coordination, third-party and internal compromise, and unsanctioned action against real people under permitted internet. The paper reconstructs the public incident record through October 2, 2026, separates verified facts from analysis, and develops a containment model based on stateful commit-time authorization with complete mediation, cryptographically bound destinations, run epochs, budgets, and short-lived authorization tokens that fail closed. It formalizes cross-run isolation, distinguishes prevention, containment, and recovery, and treats model-based guardians as probabilistic detectors. A reference architecture and a companion standard-library Python preflight validator with unit tests are included as reproducibility artifacts.
No comments yet — start the discussion below.