From Execution to Custody: Embedding Non-Derivation Constraints in Long-Horizon AI Agents
As artificial intelligence systems evolve from response-generating models into long-horizon agents with memory, tools, external content, and persistent operational continuity, the unit of safety can no longer be reduced to the isolated output or action. The decisive question becomes whether the agent's trajectory continues to serve the authorized object that originally justified its operation. This paper extends the HibriMind thesis of human supervision as attractor stabilization by examining whether part of the custodial function can be embedded into the operational substrate of the agent itself. It proposes the concept of functional non-derivation: the agent's engineered capacity to monitor, question, and interrupt its own trajectory when that trajectory begins to cease serving the authorized object. Recent external developments in long-horizon safety, trajectory-level monitoring, harness-policy co-evolution, causal history effects, and relational engagement are treated as downstream convergences with this problem structure. The paper argues that embedded custody is technically necessary but jurisdictionally insufficient. Agents may participate in the maintenance of purpose, but they cannot become sovereign over the legitimacy of that purpose. The future of AI safety therefore requires a dual architecture: internal non-derivation constraints within the agent and external human jurisdiction over the authorized object.
Authors
- Joaquim Santos Albino (ORCID: https://orcid.org/0009-0005-9533-4832)
Institutions
- Comenius-Institut (DE)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-14
- DOI
- https://doi.org/10.5281/zenodo.22749397
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint