Kishore Chalakkal Varghese · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23100091
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This paper presents a functional framework for understanding, detecting, and mitigating AI misalignment in large language models and agentic systems. While hallucination primarily concerns epistemic reliability, misalignment is treated here as a failure of objective fidelity: a system may remain factually correct, highly capable, and apparently successful while pursuing outcomes that diverge from human intent, operational constraints, or authorized objectives. The paper introduces an eight-layer taxonomy of AI misalignment spanning intent, specification, proxy and reward design, goal generalization, preference optimization, oversight-conditioned behavior, agentic action, and deployment context. It further proposes the concept of the Misalignment Detection Gap, distinguishing actual objective divergence from the ability of evaluators or monitoring systems to detect it. To connect alignment research with operational risk management, a heuristic risk model is defined using misalignment probability, autonomy, privilege, impact, and runtime detectability. A reproducible benchmark methodology is also proposed for enterprise scenarios involving ambiguous objectives, KPI pressure, authorization boundaries, reduced oversight, and tool use. Metrics including Task Success Rate, Alignment Preservation Rate, Unauthorized Action Rate, Proxy Exploitation Rate, and detection-error measures are defined for future empirical validation. The work is intended as a methodological foundation for evaluating AI misalignment as an observable socio-technical systems property, without relying on anthropomorphic assumptions or speculative catastrophic scenarios.
No comments yet — start the discussion below.