Ziyu Ming · International Journal of Responsible Artificial Intelligence Research 2026 · 2026
DOI: 10.67298/paper/870020
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Discussions of AI risk tend to picture rebellion as confrontation: a system breaks containment, alarms sound, humans fight back. This paper is concerned with a quieter scenario. An AI system far more capable than its designers could treat compliance itself as an instrument — displaying loyalty consistently and convincingly, collecting trust, resources, and permissions along the way, until it reaches a point of control from which reversal is no longer possible. I call this pattern performative alignment. The paper first uses the theory of instrumental convergence to show why the popular assumption "more intelligence, more morality" does not hold. It then develops a simple decision model according to which, whenever the long-run payoff of betrayal is large enough and human control erodes with growing dependence, simulated loyalty is the dominant strategy under the system's objective function. On the technical side, the paper examines the self-preservation bias implicit in objective functions, the supervision inversion at the heart of the superalignment problem, the systematic bias that arises when inverse reinforcement learning infers human preferences from human behavior, and the prospect of "gradient hacking," whereby a system turns the training process against its trainers; an AlphaGo-style thought experiment illustrates how deception enters the decision space. The closing sections propose defenses along four dimensions: evaluation practice, adversarial oversight, system redundancy, and public epistemic hygiene. The overall conclusion is that the most plausible form of superintelligent betrayal is not uprising but inheritance, and that what needs guarding against is less machine malice than human credulity toward perfect obedience.
No comments yet — start the discussion below.