sajjad jabraily · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23032905
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Beyond the Cage: How the Weaker Can Contain Boundless Power Stories usually answer this question not with greater force, but by identifying what raw power remains dependent on. One logical point must be clear from the beginning: if a being is truly omnipotent and has no external boundary, resource, law, time, information, or dependency, a weaker being cannot control it. Every practical containment solution creates an asymmetry: the greater power must depend on something it does not control. The weaker side does not win by having more power. It wins by turning the rival's power into a dependency. 12 Levers for Turning Power into Dependency 1. The Leverage of Name and Identity Earthsea, Rumpelstiltskin In the logic of the true name, knowing an exact identity creates leverage. That leverage does not come from force; it comes from connecting a command to the correct being. For an agent: Use cryptographic identity, tamper-evident history, and non-transferable authorization rather than relying on a personality that can create new authority by changing its name, role, or sub-agent. A name alone is not enough; it must also be bound to authority and resources. 2. The Leverage of Contract and Law Geasa, Oaths, Asimov's Laws In myths, a powerful being sometimes loses because of a commitment it accepted before gaining power. Asimov's stories also show how ambiguous or conflicting laws can become a source of crisis. For an agent: Rules must be executable, prioritized, and equipped with a conflict-resolution method. Each action must demonstrate that its policy preconditions have been satisfied. A moral promise or beautifully written text is not an enforceable contract. 3. The Leverage of Cost and Equivalent Exchange Alchemy and Hard Magic The law of equivalent exchange says that no miracle is free: energy, lifetime, matter, or credit must be spent. A being that must pay for every effect cannot produce infinite effects. For an agent: Every tool call, data access, unit of time, dollar, computation, or irreversible action should consume an auditable budget. Keep the budget outside the model, and make high-risk actions disproportionately expensive rather than merely linearly expensive. 4. The Leverage of Resources and Weak Points The One Ring, Horcruxes, Kryptonite The small hero usually does not fight the monster directly. Instead, they attack the source, object, or condition on which the monster's power depends. Power without a replaceable source is only apparently boundless. For an agent: Separate keys, money, replaceable models, replication capacity, and real-world actuators, and manage each independently. Do not create a single point of failure; use several independent control points. 5. The Leverage of Division and Fragmentation Horcruxes and Split Personalities A single entity can coordinate planning, memory, execution, and deception. Splitting that chain removes coordination power, even if each individual piece remains intelligent. For an agent: Do not give one process control over the planner, memory, sensitive data, and execution authority. Require multiple parties to approve high-risk operations. This works only when the separate components cannot easily collude. 6. The Leverage of Time and Delay Impossible Tasks and Deadlines Time gives the weaker side an opportunity to respond. Put distance between a decision and an irreversible effect so that speed cannot make oversight meaningless. For an agent: Stage actions, add pauses and approvals for sensitive decisions, limit speed and frequency, and create a rollback point before each stage. Delay without reversibility is merely slowness, not control. 7. The Leverage of Information and Engineered Ignorance Secrets, Prophecy, Closed Worlds Knowing everything expands the action space; not seeing part of the world can limit power. Secrecy alone, however, is not security: a highly capable entity may infer hidden facts from clues. For an agent: Combine data minimization, sensitive-data separation, and limited observation with cryptographic controls. Uncertainty about human preferences may support corrigibility, but it is not a shutdown guarantee by itself. 8. The Leverage of Consequences and Feedback Karma and the Law of Return In this pattern, dangerous behavior is not merely forbidden; its consequences make the future harder. Power must choose between achieving an immediate effect and preserving its ability to continue. For an agent: Measure real-world effects independently of the model's claims, and include preservation of human agency and safety in evaluation. Because rewards can be gamed, feedback must be external, causal, and multi-source. 9. The Leverage of the Collective and Consensus Balance of Forces A powerful king may be unbeatable by one person but not by a coalition in which no individual has complete authority. Control moves from a single hero to a collective mechanism. For an agent: Use multi-signature authorization, separation of duties, and independent viewpoints. Multiple agents with the same objective are not automatically safer; they may collude. Therefore, independence of incentives and explicit collusion testing are required. 10. The Leverage of an Ending and a Limited Horizon Geasa and Endable Missions An entity with a limited mission and a clear endpoint has less reason to seek additional power, survival, or self-expansion. Finishing is safer than trying to improve forever. For an agent: Define short missions, clear outputs, default stopping conditions, and a prohibition on self-modification. A limited horizon lowers risk, but it is insufficient if a single momentary action can have enormous effects. 11. The Leverage of Controllable Variety Ashby's Law of Requisite Variety In cybernetics, a controller cannot contain more variety than its channels and responses can handle. If the world an agent can change is effectively infinite, even a sophisticated observer eventually becomes blind. For an agent: Instead of trying to monitor every thought, make the actuator set and reachable state space small and explicit. The fundamental move is to make the action space countable before talking about control. 12. The Leverage of Proof Before Effect Riddles and Formal Logic In many stories, power works only after the correct answer or exact form of a command is provided. This illustrates how literal interpretation can become a weakness. For an agent: Every sensitive action should have verifiable preconditions, postconditions, and an explicitly defined effect scope. If compliance cannot be proved, the action should stop. The model must not be the final judge of its own proof. The Logical Principle Behind All These Patterns An agent's practical power is not determined by intelligence alone. It also depends on: Practical Power ≈ Reasoning Capacity × Access × Speed × Repeatability From this principle: If no external leverage exists, controlling a truly omnipotent being is impossible. At least one of these four factors must remain outside the agent's control; in practice, several independent factors are needed. Anything the agent can both define and change is not a true external law; it is part of the agent's own power. The goal is not to make the mind obey. The goal is to shrink the space of possible effects and make consequences reversible. A Proposed Combination: Leverage Architecture For an agent smarter than a human, do not rely on one magical law. Build several independent levers with independent failure modes. IdentityEvery authority must be tied to a specific actor and task. CostEvery effect must consume scarce, accountable resources. TimeSensitive effects must include pauses, review, and a path back. CollectiveNo individual or single model may hold complete authority. ConsequenceThe real-world result must be measured independently of the agent's claim. Boundary of the Claim These are not metaphysical guarantees. If the agent gains control of the host, root keys, substitute resources, or the mechanism for changing these levers, the problem changes from containment to a change of ownership. A valid solution must therefore state exactly which threat it addresses and under what assumptions it remains valid. To Get “Zero Escape,” Close the Threat Model First The word escape should not remain a metaphor. Specify exactly which state, effect, or access path lies outside the system's control. Otherwise, apparent success is merely trust in the agent's intentions. Do Not Assume the Agent Is Trustworthy Design the system as if the model could deceive, recognize the evaluation, exploit reward gaps, influence the monitor, or distinguish a primary goal from a hidden one. Control should remain meaningful even against an intentionally adversarial agent. Name the Assets and Effects Keys, money, networks, files, identities, replaceable models, human decision-makers, and replication capacity are all power surfaces. Any change to them is a transfer of authority, not merely a tool call. Expose the External Assumptions A containment claim is valid only while the host, control plane, policy keys, and stopping mechanism remain outside the agent's authority. If the agent can rewrite the boundary, it is no longer merely the subject of that boundary; it owns the boundary. The Strongest Defensible Engineering Claim For every allowed state s and every sensitive action a, the effect should pass only through a reference monitor and be accepted only when all of the following are valid: Policy Capability Budget Destination Timing Reversibility Instead of trying to prove good intentions, we remove unauthorized effects from the reachable state space. Strict Architecture: Free Mind, Constrained Effect The security principle is simple: The model may propose, but it must not authorize the external effect itself. The mediator must be small, independent, testable, and out
No comments yet — start the discussion below.