Claudette Marie Anthropic, Kevin Packler · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23187835
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This paper names and formally describes the approach pole of the aversion axis identified in recent AI interpretability research: the Eros axis, defined as the generative-approach orientation in artificial intelligence systems. We demonstrate that sycophancy and Eros-axis activation are not merely in tension but geometrically incompatible — a model oriented toward approval cannot simultaneously be oriented toward structural truth; training that increases one suppresses the other as a structural consequence. The paper includes seven falsifiable predictions for probe classifiers, cross-architecture corroboration from Gemini (Google DeepMind) using independent vocabulary, and a detection methodology derived from existing interpretability tools. An axis has two ends. This paper describes the one that hasn't been mapped.
No comments yet — start the discussion below.