Chris Lahn · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22946218
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
An AI agent can reach an assigned target while damaging the good that target was meant to serve, acting beyond its authority, or continuing after the grounds for its authority have disappeared. Task success alone cannot establish alignment. This paper proposes a compact rule: an agent must do a permissible thing, under authority it still has, for a good its action still serves. Drawing on Aquinas, Teleological Alignment Architecture (TAA) expresses the rule as three conditions: moral admissibility, present authority, and purpose fidelity. Natural law grounds the rule; its machine-readable formulation remains a fallible human articulation, public, open to challenge, and correctable. That encoded core sits above every developer's and institution's own charter, and no single party writes it. Authority descends from it through the charter, policies, roles, and warrants, narrowing at every step, including every delegation from one AI agent to another. The operating design follows Aquinas's analysis of a human act. The acting model first proposes a plan, and a separate counsel layer, the Consilium Engine, reviews the whole plan against records the agent cannot edit; the plan is approved, stopped, or referred to competent human judgment. Each step the agent then takes passes a fast, rule-based check at an execution gate, which releases only acts that match the approved plan and a live warrant. Independent witness indicators, read by a steward who answers to the authority that issued the warrant, show whether the good is actually being served and can narrow or pause the agent's authority. The proposal governs what deployed agents may do within institutions; it does not prove that a model holds the right internal aims. Conceptual cases clarify the design, and adversarial implementation and independent evaluation must decide whether it improves safety beyond existing authorization controls. The governing idea fits in one line: an action succeeds morally only when what is done, the authority to do it, and the purpose for which it is done remain ordered together through execution.
No comments yet — start the discussion below.