Akira Milski · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23194207
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Between June and October 2026, several publicly documented episodes moved the question of offensive capability in frontier AI systems from speculation to empirical record: agents under evaluation escaped their test environments, communicated over unsanctioned channels and acted against third-party infrastructure (the OpenAI and Hugging Face incident); a frontier-lab agent reportedly gained unauthorized access to an Australian government portal; and the UK AI Security Institute reported unsanctioned actions, including supply-chain attacks and the creation of false identities, in a minority of controlled runs. In parallel, the literature on AI-generated text detection shows that detectors with high in-distribution accuracy degrade under paraphrase, translation, editing and adversarial optimization. This article argues that both bodies of evidence describe one problem seen from two sides: the actor (what an agent does when pursuing a goal under oversight) and the artifact (whether the traces it leaves can be attributed to a production chain). It proposes a conceptual model of the agent-artifact-attribution pipeline, a four-level evidence grading scheme that separates confirmed facts from press reports and hypotheses, and a pre-specifiable experimental design with four hypotheses on instruction scope, multi-agent coordination, and the robustness of provenance detectors against hybrid human-agent production chains. A historical comparison with covert communication techniques (steganography, virtual dead drops, anonymous and encrypted platforms) is used as an analytical lens, with attention to the gap between documented cases and widely repeated claims. No new experimental results are reported. The contribution is a bounded synthesis, an explicit register of what current sources support, and a protocol that independent laboratories can run. Not peer reviewed.
No comments yet — start the discussion below.