Narnaiezzsshaa Truong · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22723459
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
AI systems increasingly operate in settings where organizations must demonstrate responsible behavior, policy conformance, and effective control. In response, vendors market “runtime AI governance” products that promise real-time oversight, continuous monitoring, and automated enforcement. Many such products operate at boundaries surrounding model inference: they admit or reject inputs, constrain retrieved context, filter or interrupt outputs, authorize tool calls, and record operational telemetry. These controls can be useful. Some are genuinely consequential. Their evidentiary limits, however, are frequently obscured. A refusal record can establish that a refusal occurred. A classifier output can establish that a classifier produced a result. A tool-gateway record can establish that an authorization decision was recorded at that gateway. Telemetry can establish that a telemetry system recorded or emitted a measurement. None of these artifacts, standing alone, establishes correct policy enforcement, causal policy conformance, fairness, safety, alignment, compliance, or effective control across an AI-mediated decision process. This paper presents a falsifiable evaluation framework for runtime-governance claims. It defines control surfaces, distinguishes observation from intervention and proof, identifies disclosure requirements, and specifies the evidentiary limits of common governance artifacts. It does not prescribe a proprietary implementation, data model, invariant set, scoring method, or decision rule. It defines a public standard against which runtime-governance claims should be assessed.
No comments yet — start the discussion below.