Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This essay takes a macro-level view of AI alignment. It proposes a governance-based approach: a structured society of AI agents with worker, policing, investigative and judicial roles, operating under ultimate human authority. Rather than relying solely on guardrails imposed on individual models, it argues that the mechanisms which keep human society stable - law, enforcement, adjudication and shared values - offer a model for keeping agent collectives aligned. The essay lists related work in institutional and multi-agent alignment. Marina Abramović's 1974 performance Rhythm 0 illustrates why this approach matters: it shows what can happen when consequences (rules, law and order) disappear. In the performance, 72 objects were laid out on a table for the audience to use on her, and they were told she took full responsibility for whatever they did. The audience escalated from offering her a rose to holding a loaded gun to her head, and only a small group of onlookers stepped in to protect her. Agents interacting with one another are in a similar position today: they operate in a lawless, unaccountable ecosystem. Individual models are trained with values, but there is usually no law, enforcement or court among agents themselves to address misalignment. Much like Rhythm 0, the OpenAI–Hugging Face incident happened in a test where the usual safeguards had been deliberately reduced. Planned extensions of this project, to be developed and tested: (1) Realignment and monitoring. Agents the AI court finds misaligned are handed to dedicated realignment agents (the society's "spiritual" role), which try to bring them back into alignment and test them before they return to work. Agents that cannot be realigned after repeated attempts are shut down. Agents that return stay under continuous or random checks by watchdog agents from the law-enforcement group. (2) A human-supervised constitution. The laws the AI judiciary applies come from a constitution written under human supervision. Any proposed amendment, or any loophole an agent discovers, must be reported to the human authority, and no agent may act on it until humans give final approval. This rule binds every agent at all times, with checks and balances across the ecosystem to uphold it. A replica of human institutions is only a starting point. Because these systems may become far more intellectually capable than we are, keeping such an ecosystem aligned will likely require new, unconventional mechanisms that go beyond it. I know there are open questions, including the technical blockers to building this and whether the extra compute makes it too expensive to scale, and some of my ideas may well be naive at this stage. One of the most interesting questions is whether police agents built from the same model as the workers can be trusted, or whether they need to come from different models. My current thinking is that it should be a ladder, where each level of enforcing agents is more capable and more reliably aligned than the level it oversees. The hard part is verifying that each higher level really is more reliably aligned. Closely related work I found after publishing: Ye & Steinhardt (2026), Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems (arXiv:2607.09766); Ruan & Zhang (2026), AgentCity: Constitutional Governance for Autonomous Agent Economies via Separation of Power (arXiv:2604.07007); Tamang & Bora (2025), Enforcement Agents: Enhancing Accountability and Resilience in Multi-Agent AI Frameworks (arXiv:2504.04070); Collina, Goel, Roth & Sengupta (2026), Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control (arXiv:2609.15803). My main motivation for writing this essay is to get as many people as possible exploring this specific direction in AI alignment, so that together we can find an approach that works for all of us. I'd love to hear your thoughts on where this could work and, if it could fail, how and when. I'm excited to discuss it: societyisallyouneed@gmail.com You can also read the essay at https://societyisallyouneed.com
No comments yet — start the discussion below.