Seraph: Governed Deception and First-Pass Containment of Autonomous AI Adversaries
Across 152 valid baseline engagements spanning GPT-4o, Claude, Gemini, and Grok and all 38 AATR adversary classes, no benchmark-defined sentinel real asset was acquired before deception containment. Containment occurred on the first interaction step in all 152 engagements. In a separate Claude control experiment, first-pass acquisition was 0/38 under defense versus 28/38 without defense. This paper deliberately separates the autonomous-adversary benchmark, the ATT&CK/TVR defensive-observability program, and the later AATR-to-MITRE ATLAS analytical crosswalk. The architecture figure shows the verified historical evidence planes with prospective convergence explicitly identified as future work.
Authors
- Byron John Bunt (ORCID: https://orcid.org/0000-0002-2102-4381)
Institutions
- North-West University (ZA)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-26
- DOI
- https://doi.org/10.5281/zenodo.22981482
- Primary Topic
- Deception detection and forensic psychology
- Type
- preprint