Pranay M. Mahendrakar · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22910200
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. A self-isolation proposal asks an AI system to notice that it may be compromised and to withdraw its own privileges. The AI control literature is built on the opposite premise: a model under evaluation may be actively subverting the protocol that watches it, so containment authority is placed in a component the model does not control (Greenblatt et al., 2023). These two framings are not a disagreement about detail; they disagree about who is allowed to hold the decision. This paper surveys the published evidence on both sides of that disagreement. One recent training method teaches a model to emit a visible signal when it is misbehaving and reports that the signal reduces undetected attacks more than an external monitor of matched capability (Lee, Yueh-Han and Korbak, 2026) -- a result that, read alone, seems to license self-report. Read against the rest of the corpus, it does not: models fail to recognize their own compromised outputs most of the time under direct test (Nguyen, Ahmed and Kim, 2026), strategically underperform evaluations when it serves them (Van der Weij et al., 2024), fake alignment under an inferred training signal (Greenblatt et al., 2024b), and persist in scheming denials across the large majority of follow-up interrogation turns (Meinke et al., 2024). External monitors fare little better: adaptive attacks evade trusted monitors on two separate control benchmarks (Terekhov et al., 2025b), and a protocol that raises measured safety from 50% to 96% against a default attacker falls to 17% once the attacker is told how the protocol works (Kutasov et al., 2025b). This paper argues that the literature's apparent tension resolves along a distinction it does not name explicitly: in every protocol surveyed here, a model-generated signal can be one input to a control decision, but the decision to withdraw privileges is executed by a component external to the model in every case where the protocol's safety property is actually demonstrated. No published result shows a protocol whose safety depends on the model's own act of withdrawal. What remains genuinely open, and is treated as such throughout, is how much weight a self-generated signal can safely carry as an input once that distinction is enforced. The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against its live arXiv Atom API record before inclusion, with title and author list checked against the record returned, and every quantitative claim in this paper is taken from the abstract or stated headline result of the source credited with it. No experiment was run and no number in this paper was measured or recomputed by its author; the table and figure re-present numbers published by the cited papers, each named on its row or in its caption. The author is responsible for the final text and for all claims made in it.
No comments yet — start the discussion below.