Maria Luz Madariaga · · 2026
DOI: 10.33774/coe-2026-wp8dr
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Human oversight is the control on which most AI-agent governance rests, yet deployments rarely record whether it was exercised — only that a mechanism exists. This paper treats that gap as a documentation-and-control failure and names its acute form override theatre: a single actor clears a policy-violating agent action with one dis-position, holding the initiate–authorise–record chain end to end, so the appearance of oversight substitutes for the substance. Drawing from financial control practice ra-ther than machine learning, the paper makes one claim: the human override of a policy-violating agent action is a governed decision that must be segregated, leased, scoped, version-locked, and recorded on an independent evidence trail. Single-tier oversight, lacking these proper-ties, cannot be audited and cannot be trusted. The deci-sion is situated in the three lines of defence: first line disposes under segregation of duties; second line moni-tors the override path and raises versioned policy and routing changes with the business and legal; third line audits independently and does not gate release. The loop improves the control, not the model — recorded disposi-tions are never written back as model updates. A non-destructive proof-of-concept runs three conditions on the same eight items: unsegregated two-tier (three viola-tions sent to override and disposed by the same actor, SoD flag true on every override); segregated two-tier (four violations sent to override and upheld by an inde-pendent authoriser, SoD flag false); and single-tier (two of four violations approved in one click). The empirical claim is fea-sibility and face validity.
No comments yet — start the discussion below.