Gregorio von Hildebrand · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22834587
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
When an AI agent acts destructively, who notices, and can a third party later check that anything was watching? We read sixteen documented incidents from July 2025 to September 2026 and find that the first detector was almost always the harmed person, that detection took minutes when a person was present and days when not, and that no incident had an automated third-party alarm. We then survey 38 deployed controls and nine standards and find an empty intersection: the controls that block an action before it runs report only to the operator, and the things that reach beyond the operator do not block. We contribute no new policy engine. We show that a deliberately simple deterministic gate becomes third-party-verifiable evidence that oversight existed and fired, once every decision and the gate's own presence are written to a hash-chained ledger and its head is sealed into a public transparency log and timestamped by an authority that neither the operator nor the gate's author controls. We evaluate the sentinel by replaying it over our own agent fleet's complete git history (364 commits) and over a corpus of 50 destructive commands drawn from the documented incidents, 40 benign look-alikes, 36 variant forms and 200 instruction payloads. It catches 50 of 50 destructive commands with no false positive, 34 of 36 variants, and no decision is changed by any payload, because nothing in the blocking path reads text as instruction. The replay found one contradiction between an agent's standing order and its declared scope. The sentinel runs on the fleet that wrote this paper; its record is public and its first attestation is cited in the text. Disclosure: this work was researched, built and drafted with Vigilia, an autonomous AI system operated by Dear Wise Earth Costa Rica SRL, under the direction of the human author, who reviewed the claims and takes responsibility for them. Code, ledger, attestations and evaluation data: https://github.com/GvHildebrand/sentinel-hook. Rendered version: https://aivigilia.com/papers/vigilia-sentinel-2026.pdf
No comments yet — start the discussion below.