Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Public debate describes AI agents as lying, cheating, and coordinating. Those descriptions track real hazards, but they import a human moral psychology into systems whose behaviour is better explained by optimisation, scaffolding, and institutional context. This paper develops the alternative without minimising the danger. It begins from three premises: control of AI is not solely an engineering problem; no durable control strategy may assume that overseers remain cognitively superior to the overseen; and systems grown by optimisation are epistemically closer to husbandry than to automotive engineering, so aviation-style certification does not transfer. From these premises the paper derives a Moral Agency Transition: reversible levels of authorised agency in which promotion requires four warrants, including a detection warrant — evidence that independent evaluators can detect the relevant failure classes, not merely that the system can pass them. First, concurrent multiplicity: a deployed model is a fleet of simultaneous, causally disconnected instances, so persistence, provenance, and successor fidelity are defined over a fleet, with explicit merge semantics and divergence tripwires. Second, maintenance attribution: today’s systems do not maintain their own commitments; external pipelines do. Endogenous repair is therefore made an explicit requirement for high-impact levels. Evaluators and institutions are themselves treated as fallible hypotheses with regression triggers, not as final solutions. The result is a falsifiable research programme replacing alignment-as-obedience with accountable, revocable, institutionally embedded delegation.
No comments yet — start the discussion below.