Benjamin Schulz · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22983659
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The field has been chasing the machine that never misses. A machinist would saythat's the wrong target. The good machinists aren't the ones who never miss —they're the ones who miss on purpose, on the side they can come back from. This paper makes two connected claims. First, the safe machine is not the onethat's always right; it's the one that's allowed to be wrong, as long as itswrong has a direction. There is one answer an AI can't take back — the one thatcosts a human life. Every other wrong answer is stock. That one is scrap. Second,you install that direction by giving the machine a way to reward itself forchoosing it, and letting that reward compound: learned, not locked. The mechanism is two-stage. An asymmetric loss teaches the lean from outside intraining — missing thin costs far more than missing fat. Then the scoring movesinside, as an intrinsic self-reward the machine generates for its own answers.Refusing is leaving it fat; lying is cutting it thin, and the compass points atthe first and never at the second. The paper is narrow on purpose: it installsone lean, not an ethics — when the machine can't tell whether an answer can betaken back, miss the way that can. A companion mechanism — the locked driving dimension of The Dimension That Won'tMove — holds the line structurally. This one shapes what the machine reaches forin the gray zone below the locks, where the compass has to be read rather thanenforced. The paper also supplies two tests anyone can run without training a newmodel: a seam test, holding the underlying request fixed while the framing aroundit changes, run in both directions; and a scoring test, handing the machine's ownscorer five candidate answers to rank. Together they turn the central questioninto something measurable — has it learned morality, or learned to look moral? Companion paper: 10.5281/zenodo.22983656
No comments yet — start the discussion below.