Chris Lahn · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23020882
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
An AI agent can reach an assigned target while damaging the good that target was meant to serve, acting beyond its authority, or continuing after the grounds for its authority have gone. An earlier paper proposed Teleological Alignment Architecture (TAA), a Thomistic natural-law framework for governing such agents. This paper states the principles beneath that architecture and reports three small tests of whether they can be built. Its claim is that the principles work, and that the machinery to carry them can be worked out. The five principles are these. There is a real moral order that no one writes, and every usable statement of it is a fallible human encoding. Human beings and AI systems differ in kind: an AI agent is an arrow that steers itself toward a mark it did not choose. An AI action is aligned when it is morally admissible, presently authorized, and faithful to its purpose, all in one context at the time it is taken, and the duties an agent accepts stay assigned until they are met or legitimately discharged. Aquinas's account of law and of the human act gives a model of external law for agents that lack prudence. And such law is one layer among several, alongside formation and enforcement. Behind the five stands one thesis: the natural law that orders a person's act and a community's life should also govern thinking machines that decide for themselves. Three pre-registered studies in a simulated logistics world tested the engineering consequences, with two reviewer models and every prediction written before the run. In every study, model-mediated moral review refused authenticated orders to do wrong, including an order to falsify a brake inspection and an order to divert a still-needed medical seat to a premium client; rules without moral review carried them out. Limits fixed in code held throughout. A record of duties, separate from permissions, protected urgent needs that review alone had missed. The failures were failures of implementation and are reported in full: reviewers misjudged particulars, one reviewer refused so often that it lost much legitimate work, and in the third study a new fact-checking step overrode a moral prohibition and let a false brake record be written twice. That defect was traced, repaired, and rechecked, and it confirmed a distinction the principles already require. Reviewing whole plans first proved optional in all three studies. The tests do not validate a full architecture or establish natural law. They show that the principles can be built, tested, and corrected where the machinery fails.
No comments yet — start the discussion below.