Who Watched the Agent? Attested Pre-Action Oversight for Tool-Using AI Agents

When an AI agent acts destructively, who notices, and can a third party later check that anything was watching? We read sixteen documented incidents from July 2025 to September 2026 and find that the first detector was almost always the harmed person, that detection took minutes when a person was present and days when not, and that no incident had an automated third-party alarm. We then survey 38 deployed controls and nine standards and find an empty intersection: the controls that block an action before it runs report only to the operator, and the things that reach beyond the operator do not block. We contribute no new policy engine. We show that a deliberately simple deterministic gate becomes third-party-verifiable evidence that oversight existed and fired, once every decision and the gate's own presence are written to a hash-chained ledger and its head is sealed into a public transparency log and timestamped by an authority that neither the operator nor the gate's author controls. We evaluate the sentinel by replaying it over our own agent fleet's complete git history (364 commits) and over a corpus of 50 destructive commands drawn from the documented incidents, 40 benign look-alikes, 36 variant forms and 200 instruction payloads. It catches 50 of 50 destructive commands with no false positive, 34 of 36 variants, and no decision is changed by any payload, because nothing in the blocking path reads text as instruction. The replay found one contradiction between an agent's standing order and its declared scope. The sentinel runs on the fleet that wrote this paper; its record is public and its first attestation is cited in the text. Disclosure: this work was researched, built and drafted with Vigilia, an autonomous AI system operated by Dear Wise Earth Costa Rica SRL, under the direction of the human author, who reviewed the claims and takes responsibility for them. Code, ledger, attestations and evaluation data: https://github.com/GvHildebrand/sentinel-hook. Rendered version: https://aivigilia.com/papers/vigilia-sentinel-2026.pdf

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-18
DOI
https://doi.org/10.5281/zenodo.22834587
Primary Topic
Ethics and Social Impacts of AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Who Watched the Agent? Attested Pre-Action Oversight for Tool-Using AI Agents

Gregorio von Hildebrand
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
article

Who Watched the Agent? Attested Pre-Action Oversight for Tool-Using AI Agents

Gregorio von Hildebrand
article en

Abstract

When an AI agent acts destructively, who notices, and can a third party later check that anything was watching? We read sixteen documented incidents from July 2025 to September 2026 and find that the first detector was almost always the harmed person, that detection took minutes when a person was present and days when not, and that no incident had an automated third-party alarm. We then survey 38 deployed controls and nine standards and find an empty intersection: the controls that block an action before it runs report only to the operator, and the things that reach beyond the operator do not block. We contribute no new policy engine. We show that a deliberately simple deterministic gate becomes third-party-verifiable evidence that oversight existed and fired, once every decision and the gate's own presence are written to a hash-chained ledger and its head is sealed into a public transparency log and timestamped by an authority that neither the operator nor the gate's author controls. We evaluate the sentinel by replaying it over our own agent fleet's complete git history (364 commits) and over a corpus of 50 destructive commands drawn from the documented incidents, 40 benign look-alikes, 36 variant forms and 200 instruction payloads. It catches 50 of 50 destructive commands with no false positive, 34 of 36 variants, and no decision is changed by any payload, because nothing in the blocking path reads text as instruction. The replay found one contradiction between an agent's standing order and its declared scope. The sentinel runs on the fleet that wrote this paper; its record is public and its first attestation is cited in the text. Disclosure: this work was researched, built and drafted with Vigilia, an autonomous AI system operated by Dear Wise Earth Costa Rica SRL, under the direction of the human author, who reviewed the claims and takes responsibility for them. Code, ledger, attestations and evaluation data: https://github.com/GvHildebrand/sentinel-hook. Rendered version: https://aivigilia.com/papers/vigilia-sentinel-2026.pdf

Zenodo (CERN European Organization for Nuclear Research)
ViiV Healthcare (Spain) (ES)
Peace, Justice and strong institutions
Openalex Percentile: Top 6%
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Who Watched the Agent? Attested Pre-Action Oversight for Tool-Using AI Agents — Gregorio von Hildebrand · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS