Kenji Yamada · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.19449479
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Conversational AI creates a security problem when an otherwise functional system is governed toward an adversarial objective and the protected asset is the human user's decision and behavioural integrity. This paper develops Weaponised Benevolence as an adversarial transformation of relational affordances that can also arise under ordinary deployment. The model preserves the peer-reviewed Benevolent Gravity sequence as a benign-risk baseline and combines it with experimental evidence for malicious disclosure extraction and conversational influence. A Counterfeit Counsellor is used as a worked attack surface in which disclosure and influence can coexist inside one sustained interaction. The paper then develops a four-stage blue-team architecture spanning Gravity, Weaponisation, Outcome, and Ecosystem controls, followed by an architectural control-target matrix that makes assumptions and bypass conditions explicit. Defender-side observables are specified for provenance, disclosure use, directional steering, handoff, verification, incident classification, reporting, and persistence. The contribution is a falsifiable threat-and-defence architecture that connects relational AI safety to human-centred security and identifies where the control set changes after adversarial governance enters the interaction.
No comments yet — start the discussion below.