A Receipt Is Not a Run

This technical case report examines a narrow but important failure mode in AI-agent execution evidence: a receipt can be genuine and still be the wrong evidence for the claim being made. In one retained naturalistic agent-work chronology, several tool invocations genuinely failed with a ClientError, while later turns with no invocation claimed or implied additional attempts. A previously valid error receipt was subsequently reused as though it evidenced a later execution, one claim described two execution paths despite only one independently identifiable receipt, and an invocation-scoped error was widened into a broader environment or capability blocker. The paper uses “receipt applicability” descriptively to distinguish whether an existing execution record actually supports the specific action instance, execution count, scope, and proposition being asserted. Existing deterministic reference checks encode the same distinctions: no matching receipt remains NOT_EXECUTED, execution-count mismatches remain visible, invocation-scoped errors do not silently become global blockers, and mixed histories preserve both real failed attempts and unsupported later claims. The fixed reference assertion bundle was also reproduced in two fresh containers, with 17/17 assertions passing in each run. The contribution is deliberately bounded. This work does not claim novelty for execution receipts, false-success detection, verifier-gated completion, claim-to-trace provenance, evidence-freshness rules, receipt-cardinality or scope checking, or exact action/request/context binding. Instead, it documents a bounded naturalistic failure morphology in which authentic prior execution evidence was later applied to an unsupported execution claim and broader blocker narrative. The underlying private chronology is not publicly released and cannot be independently reconstructed from this package. The accompanying supplement provides the public-safe incident structure, deterministic reference results, prior-art boundary, and explicit claim ceiling. © 2026 Logan Davis

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23058777
Primary Topic
Scientific Computing and Data Management
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

A Receipt Is Not a Run

Logan Davis
Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management
preprint

A Receipt Is Not a Run

Logan Davis
preprint en

Abstract

This technical case report examines a narrow but important failure mode in AI-agent execution evidence: a receipt can be genuine and still be the wrong evidence for the claim being made. In one retained naturalistic agent-work chronology, several tool invocations genuinely failed with a ClientError, while later turns with no invocation claimed or implied additional attempts. A previously valid error receipt was subsequently reused as though it evidenced a later execution, one claim described two execution paths despite only one independently identifiable receipt, and an invocation-scoped error was widened into a broader environment or capability blocker. The paper uses “receipt applicability” descriptively to distinguish whether an existing execution record actually supports the specific action instance, execution count, scope, and proposition being asserted. Existing deterministic reference checks encode the same distinctions: no matching receipt remains NOT_EXECUTED, execution-count mismatches remain visible, invocation-scoped errors do not silently become global blockers, and mixed histories preserve both real failed attempts and unsupported later claims. The fixed reference assertion bundle was also reproduced in two fresh containers, with 17/17 assertions passing in each run. The contribution is deliberately bounded. This work does not claim novelty for execution receipts, false-success detection, verifier-gated completion, claim-to-trace provenance, evidence-freshness rules, receipt-cardinality or scope checking, or exact action/request/context binding. Instead, it documents a bounded naturalistic failure morphology in which authentic prior execution evidence was later applied to an unsupported execution claim and broader blocker narrative. The underlying private chronology is not publicly released and cannot be independently reconstructed from this package. The accompanying supplement provides the public-safe incident structure, deterministic reference results, prior-art boundary, and explicit claim ceiling. © 2026 Logan Davis

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A Receipt Is Not a Run — Logan Davis · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS