False Execution in Stateful AI Workflows: An Exploratory Failure Analysis

Tool-using AI systems increasingly operate in stateful workflows in which successful task completion requires not only generating an appropriate response but also performing actions that change an external state. This exploratory failure analysis documents repeated cases from a longitudinal, stateful AI workflow in which an AI system produced a semantically complete terminal response while a required external action remained unexecuted and unevidenced. In the observed cases, the system possessed the relevant procedural knowledge and execution capability, and the omitted action could subsequently be performed. Notably, a minimal user follow-up containing little or no new task-relevant operational information was sufficient to reopen the action path and produce the previously omitted external state transition. We distinguish semantic completion, operational completion, and evidenced external state transition, and reconstruct the observed failure as a divergence between these states rather than as a simple absence of tool capability. Successful autonomous tool-use episodes from the same workflow provide contrast cases showing that the system was capable of initiating and completing external actions without user prompting under other conditions. The paper does not establish a causal mechanism. We examine several candidate explanations, including premature terminal acceptance, weak binding between required external actions and the operational definition of completion, and a possible product-identity mismatch in which a semantically salient intermediate artifact is treated as the final process product. The central empirical contribution is the documentation of a repeated gap between semantic and operational completion, including recovery after low-information user intervention.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.22309886
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

False Execution in Stateful AI Workflows: An Exploratory Failure Analysis

Alen Širola
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
article

False Execution in Stateful AI Workflows: An Exploratory Failure Analysis

Alen Širola
article en

Abstract

Tool-using AI systems increasingly operate in stateful workflows in which successful task completion requires not only generating an appropriate response but also performing actions that change an external state. This exploratory failure analysis documents repeated cases from a longitudinal, stateful AI workflow in which an AI system produced a semantically complete terminal response while a required external action remained unexecuted and unevidenced. In the observed cases, the system possessed the relevant procedural knowledge and execution capability, and the omitted action could subsequently be performed. Notably, a minimal user follow-up containing little or no new task-relevant operational information was sufficient to reopen the action path and produce the previously omitted external state transition. We distinguish semantic completion, operational completion, and evidenced external state transition, and reconstruct the observed failure as a divergence between these states rather than as a simple absence of tool capability. Successful autonomous tool-use episodes from the same workflow provide contrast cases showing that the system was capable of initiating and completing external actions without user prompting under other conditions. The paper does not establish a causal mechanism. We examine several candidate explanations, including premature terminal acceptance, weak binding between required external actions and the operational definition of completion, and a possible product-identity mismatch in which a semantically salient intermediate artifact is treated as the final process product. The central empirical contribution is the documentation of a repeated gap between semantic and operational completion, including recovery after low-information user intervention.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

False Execution in Stateful AI Workflows: An Exploratory Failure Analysis — Alen Širola · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS