Wolfgang Reinl · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22868534
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Most AI alignment frameworks evaluate whether an artificial system behaves consistently with a human objective, preference model, normative standard, or oversight process. This paper argues that a further problem emerges when the entities on both sides of the alignment relation are themselves changing. Human preferences can change and can be influenced by AI; human values need not be reducible to stable preferences; future humans or human-descended agents may differ materially from present humans; and increasingly capable AI may participate in the development of successor AI systems. The resulting problem is not merely dynamic alignment. It is alignment across a coupled transition in which humans H_t, values V_t, artificial systems A_t, and available knowledge K_t may all change, potentially on sharply different timescales. We call this problem class Transitional Alignment and organize it around three coupled problems. The Moving-Target Problem asks what long-run alignment can mean when the human side of the alignment relation is not static. The Differential-Rate Problem asks what happens when AI capability and successor creation proceed faster than meaningful human deliberation, adaptation, or governance. The Successor-Transport Problem asks whether any desirable safety property survives delegation, self-modification, automated AI research, model replacement, and successor creation. These problems remain even if a present system is well aligned. The paper rejects a tempting solution: instructing advanced AI to maximize agency, option value, pluralism, or procedural legitimacy. Such quantities are proxies and can themselves be Goodharted. Transitional Alignment therefore treats agency, corrigibility, non-manipulation, optionality, and deliberative capacity as candidate properties for adversarial evaluation, not as a scalar objective for superintelligence. Its main contribution is a research decomposition and a set of conditional claims: if meaningful participation is to persist through transformative AI, transition speed and successor-property transport must become first-class alignment variables. A worked coupled-failure case shows how small per-transition losses can accumulate across several successor cycles before one meaningful human review, creating cumulative unreviewed degradation; the broader accumulation of unresolved degradation, uncertainty, and verification gaps is termed transition debt.
No comments yet — start the discussion below.