Xiangsheng YU · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22859356
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Runtime security monitoring of AI agents has converged on behavioural signals — tool calls, execution traces, invocation sequences — yet remains empirically driven: baselines are learned from samples of benign behaviour, and no account is given of when such a monitor is necessarily effective. We argue that the difficulty is not one of detector design but of interface selection, and that most candidate interfaces are structurally unusable rather than merely weak. We further point out that all AI reasoning systems depend on an axiom library. In current AI this library exists in implicit form, embedded in training data and model weights, and its role is to anchor output when a gap appears in the logical chain. The axiom-library invocation interface is therefore an observable common to all AI, and there is no category of “AI without an axiom library”; the applicability of this framework is not limited by model form, scale, or training regime. We first establish an ontological premise: an error state of a physical process must be distinguishable from its non-error states, otherwise the error has no referent. Completeness is therefore discovered, not designed. We then show that under current architectures — which must produce output — a gap in the logical chain must be compensated, and that the vehicle of compensation is the invocation of axiom-library constants, the anchor onwhich output generation depends. Hence the trace left by an error is not merely present but located. The monitored quantity is the invocation behaviour of axiom-library constants rather than business-layer parameters, which are engineering level expressions that an adversary can manipulate. We show a monitoring port is usable only if it is (i) quantifiable in a principled way, (ii) decidable — its baseline fixed by architecture rather than by task, and (iii) complete under a stated threat model. Semantic ports fail (ii) and (iii); internal-state ports fail (i) and (ii) — we compute a discriminative signal-to-noise ratio of 0.33 for attention-based monitoring, i.e. the signal is buried in cross-task variation; resource side channels fail (iii) because they are output-side aggregates — many-to-one, non-monotone under efficiency modulation, and drifting with demand. The axiom-library invocation interface is the only external observable on the demand side, and hence the only one satisfying all three. On this interface we build a three-channel monitor (frequency, deviation, smoothness) and prove: completeness (every effective attack leaves a trace), order-invariance (under binary output with no intermediate state, aggregate output depends only on the multiset of calls — reordering flips output in 0% of trials versus 94.5% once intermediate state is permitted), non-compensability (perturbations are non-negative, so suppression on one channel yields zero escape benefit), and zero gap (with thresholds calibrated to self-correction capacity, exceeding threshold and being effective are equivalent). A change absent from all three channels is by definition not an attack — so the framework admits neither false negatives nor false positives of the “detected but harmless” kind; the autoimmune rate intrinsic to the AIS paradigm is determined by scenario calibration and reported as measured (§7.3). We locate the contribution within artificial immune systems (AIS): the monitor is a self/non-self discrimination layer that does not identify pathogens. Classical AIS relies on negative selection, whose detector count grows combinatorially with the dimension of the self-space (we compute 48 dimensions to be infeasible). By deriving the self-set analytically from architectural constraints rather than statistically from samples, the present framework removes this bottleneck: three observables, independent of problem dimension. Because the monitor is a bypass that does not participate in inference, evolution of the monitored system changes readings but not the criterion. Finally, we argue that explicitization of the axiom library is the necessary path toward verifiable, auditable, and white-box AI, and give an observability spectrum across which every tier falls within the framework’s scope.
No comments yet — start the discussion below.