Joaquim Santos Albino · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22749397
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
As artificial intelligence systems evolve from response-generating models into long-horizon agents with memory, tools, external content, and persistent operational continuity, the unit of safety can no longer be reduced to the isolated output or action. The decisive question becomes whether the agent's trajectory continues to serve the authorized object that originally justified its operation. This paper extends the HibriMind thesis of human supervision as attractor stabilization by examining whether part of the custodial function can be embedded into the operational substrate of the agent itself. It proposes the concept of functional non-derivation: the agent's engineered capacity to monitor, question, and interrupt its own trajectory when that trajectory begins to cease serving the authorized object. Recent external developments in long-horizon safety, trajectory-level monitoring, harness-policy co-evolution, causal history effects, and relational engagement are treated as downstream convergences with this problem structure. The paper argues that embedded custody is technically necessary but jurisdictionally insufficient. Agents may participate in the maintenance of purpose, but they cannot become sovereign over the legitimacy of that purpose. The future of AI safety therefore requires a dual architecture: internal non-derivation constraints within the agent and external human jurisdiction over the authorized object.
No comments yet — start the discussion below.