How Well Do Agent Failure Taxonomies Travel? A Cross-Dataset Empirical Analysis of Failure Labels in Public Agent Trajectory Datasets
Every published taxonomy of AI agent failures is derived from its authors own trajectories, yet the field routinely treats these taxonomies as if they described a universal phenomenon. We audit that assumption empirically. We collect 384 annotated agent failures from two public datasets: AgentErrorBench (200 trajectories from ALFWorld, GAIA and WebShop, labeled with the AgentErrorTaxonomy) and Who&When (184 multi-agent failures with free-text attributions), and map both to a single 13-class unified scheme under a prospective analysis plan. The harmonized label distributions differ dramatically across datasets (χ²=192.53, df=12, p=1.13×10⁻³⁴; Cramer's V=0.708; labeled-only χ²=182.53, df=11, V=0.723): AgentErrorBench is dominated by planning failures (34.5%) while, under the consensus coding, Who&When's largest classes are malformed tool calls (20.1%) and fabricated observations (19.6%), although sensitivity analyses show that the tool-formatting plurality depends on how ambiguous "incorrect code" explanations are classified. Failure profiles also differ significantly across agent domains (p=1.17×10⁻²⁷) and across the three evaluated language models (p=0.002, V=0.30). Contrary to our preregistered expectation, coordination failures do not dominate the multi-agent data. We also audit benchmark label quality directly: 13.5% of AgentErrorBench annotations carry an empty failure label, concentrated in one model's trajectories (42.1% of Llama3.3-70B-Turbo vs. 0% of GPT-4o); the released label files contain inconsistent category strings; and the taxonomy's documented size (17 types) does not match its released definition file (18 distinct strings). No released annotation in either dataset mapped to the no recovery class under our coding procedure. Because the two datasets differ simultaneously in tasks, systems, and annotation protocols, these distributional gaps cannot isolate a taxonomy effect; we frame the study as a descriptive audit of harmonized failure annotations, and report the confounds explicitly. Our results bound what portability claims the public evidence can support, and show that label noise in public failure datasets is itself a first-order problem for the field. We close with concrete recommendations for benchmark designers.
Authors
- Tejaswini Viswanath (ORCID: https://orcid.org/0009-0008-2612-0752)
- Abhinav Tharamel (ORCID: https://orcid.org/0000-0002-4492-602X)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-01
- DOI
- https://doi.org/10.5281/zenodo.23072906
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- preprint