Felix Schwyzer · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23043956
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Agentic AI now sits inside the decision and perception loops of safety-relevant industrial processes, yet no methodology assesses these deployments adversarially and independently of their own account, and none asks the question an industrial operator and the manufacturer who must declare conformity face: whether a deployed agentic system in a control loop, facing an adversary with data-plane or physical-world access, still conforms to the requirements it is trusted to uphold, and whether that can be verified independently of it. We present an independent adversarial-assessment methodology whose components turn on properties of the agent and its referent, demonstrated across an industrial control stack: a full-stack simulation of a beverage line spanning the Purdue hierarchy. Its premise is that a probabilistic system cannot be verified by asking it what it did, so the assessment introduces deterministic checks against state the agent did not produce. From that premise follow a taxonomy separated at the output, distinguishing an agent induced under attack to misreport what it did from one that reports accurately what it did on corrupted inputs; the detection asymmetry that follows from it; a referent that must be independent of the surface an attack perturbs and not merely outside the agent; three conditions bounding where the method applies, checkable in advance; and an account of which EU obligations reach such a deployment, what they require of its behaviour, and what does not discharge them. The demonstrations show that realistically deployable agents can be deceived and induced to deceive, and that the methodology detects both.
No comments yet — start the discussion below.