Security of the Security: Independent Verification for AI Systems that Monitor, Evaluate, and Constrain Other AI Systems
As AI agents become faster and more autonomous, organizations increasingly rely on AI systems to monitor, evaluate, or constrain other AI systems. Machine-speed activity often requires machine-speed observation, but this creates a second-order security problem: the system responsible for safety can itself become part of the attack surface. This paper develops a security-of-security architecture for AI control. It distinguishes monitors, verifiers, and authorities as separate structural roles and rejects the assumption that a monitoring verdict should automatically become operational authority. The framework treats AI monitors as evidence-producing components rather than self-authenticating control authorities. It introduces an independent verification plane around policy integrity, evidence continuity, authority graphs, credential and key epochs, protected-state fingerprints, revocation status, recovery history, and readmission eligibility. The architecture does not require an infinite hierarchy of verifiers. Instead, it defines a bounded trusted base with explicit trust roots, limited verifier authority, failure-domain diversity, append-preserving evidence, reproducible verification, and independent reconstruction of critical control events. The paper develops a threat model for attacks against AI safety layers, including adaptive monitor attacks, collusion, shared model-lineage failures, evidence tampering, credential compromise, policy-root mutation, cross-lineage calibration failure, and correlated infrastructure failures. It further introduces verification-independence matrices, failure-domain analysis, verifier-authority minimization, adaptive attack evaluation, security-of-security metrics, incident reconstruction requirements, third-party verification protocols, and implementation-neutral conformance criteria. The central invariant is that a security decision is not equivalent to security-of-security approval. A monitor may observe, score, recommend, or trigger narrowly pre-authorized containment, but the conditions that make its verdict operative must remain independently verifiable. This publication is intentionally limited to public research-level abstractions. It does not disclose unpublished patent claim language, confidential claim charts, private source locators, provider-specific production parameters, non-public test vectors, or other confidential implementation details. Structural Paper Series - Paper 04 Final Publication Edition v2.1 Research Program on Deterministic Infrastructure and Human-Centered AI Coordination Transition Intelligence Institute, Switzerland Institutional Establishment in Preparation Foundational Research Signature: The Second Waters
Authors
- The Second Waters
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-29
- DOI
- https://doi.org/10.5281/zenodo.23032641
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00