A Receipt Is Not a Run
This technical case report examines a narrow but important failure mode in AI-agent execution evidence: a receipt can be genuine and still be the wrong evidence for the claim being made. In one retained naturalistic agent-work chronology, several tool invocations genuinely failed with a ClientError, while later turns with no invocation claimed or implied additional attempts. A previously valid error receipt was subsequently reused as though it evidenced a later execution, one claim described two execution paths despite only one independently identifiable receipt, and an invocation-scoped error was widened into a broader environment or capability blocker. The paper uses “receipt applicability” descriptively to distinguish whether an existing execution record actually supports the specific action instance, execution count, scope, and proposition being asserted. Existing deterministic reference checks encode the same distinctions: no matching receipt remains NOT_EXECUTED, execution-count mismatches remain visible, invocation-scoped errors do not silently become global blockers, and mixed histories preserve both real failed attempts and unsupported later claims. The fixed reference assertion bundle was also reproduced in two fresh containers, with 17/17 assertions passing in each run. The contribution is deliberately bounded. This work does not claim novelty for execution receipts, false-success detection, verifier-gated completion, claim-to-trace provenance, evidence-freshness rules, receipt-cardinality or scope checking, or exact action/request/context binding. Instead, it documents a bounded naturalistic failure morphology in which authentic prior execution evidence was later applied to an unsupported execution claim and broader blocker narrative. The underlying private chronology is not publicly released and cannot be independently reconstructed from this package. The accompanying supplement provides the public-safe incident structure, deterministic reference results, prior-art boundary, and explicit claim ceiling. © 2026 Logan Davis
Authors
- Logan Davis (ORCID: https://orcid.org/0009-0006-8244-5610)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23058778
- Primary Topic
- Scientific Computing and Data Management
- Type
- preprint