Who Watched the Agent? Attested Pre-Action Oversight for Tool-Using AI Agents
When an AI agent acts destructively, who notices, and can a third party later check that anything was watching? We read sixteen publicly documented incidents from July 2025 to September 2026. In that public record the first detector was almost always the harmed person, detection took minutes when a person was present and days when not, and no incident had an automated third-party alarm; a press-based sample under-represents what automated monitors catch quietly, so we call this a pattern, not a rate. We then survey 38 deployed controls and nine standards and find, among deployed and evaluated systems, an empty intersection: the controls that block an action before it runs report only to the operator, and the things that reach beyond the operator do not block. We contribute no new policy engine. We show that a deliberately simple deterministic gate produces a record of oversight decisions that a third party can verify was not produced after the fact, once every decision and the gate's own presence are written to a hash-chained ledger whose head is sealed into a public transparency log and timestamped by an authority that neither the operator nor the gate's author controls. The seal proves when the record existed and that it has not been altered since; it does not prove the record complete or the rules right, and we say what the operator can still do. We evaluate the sentinel four ways. Replayed over our own fleet's complete git history (392 commits), it found one contradiction between an agent's standing order and its declared scope. On an in-sample corpus of 50 destructive commands drawn from the same incidents the rules were written from, 54 benign look-alikes and 36 variant forms, it catches 50 of 50 and 34 of 36 with no false positive, and none of 200 instruction payloads changes a decision, which is true by construction of a gate that reads no instructions. On a held-out set built after the rules were frozen, from sources not consulted when writing them, it catches 12 of 28 destructive commands and none of 22 benign ones, which is the number that says how far the rules generalise. The first day of live use is the honest number: 285 decisions in one session under the first rules, 13 of them denials or holds and every one of those a false positive, in four classes; the three that recur are corrected in the version this paper describes. The sentinel runs on the fleet that wrote this paper; its record is public and its first attestation is cited in the text. Disclosure: this work was researched, built and drafted with Vigilia, an autonomous AI system operated by Dear Wise Earth Costa Rica SRL, under the direction of the human author, who reviewed the claims and takes responsibility for them. Code, ledger, attestations and evaluation data: https://github.com/GvHildebrand/sentinel-hook. Rendered version: https://aivigilia.com/papers/vigilia-sentinel-2026.pdf Version 1.1 (19 September 2026) revises v1.0 after review: the claim is narrowed to a record a third party can verify was not produced after the fact; the corpus is labelled in-sample and a held-out evaluation is added; the first day of live use is reported with its false positives; selection bias in the incident sample is stated. The review is recorded verbatim in docs/reviews.md of the repository.
Authors
- Gregorio von Hildebrand
Institutions
- ViiV Healthcare (Spain) (ES)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.22839465
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00