How Well Do Agent Failure Taxonomies Travel? A Cross-Dataset Empirical Analysis of Failure Labels in Public Agent Trajectory Datasets

Every published taxonomy of AI agent failures is derived from its authors own trajectories, yet the field routinely treats these taxonomies as if they described a universal phenomenon. We audit that assumption empirically. We collect 384 annotated agent failures from two public datasets: AgentErrorBench (200 trajectories from ALFWorld, GAIA and WebShop, labeled with the AgentErrorTaxonomy) and Who&When (184 multi-agent failures with free-text attributions), and map both to a single 13-class unified scheme under a prospective analysis plan. The harmonized label distributions differ dramatically across datasets (χ²=192.53, df=12, p=1.13×10⁻³⁴; Cramer's V=0.708; labeled-only χ²=182.53, df=11, V=0.723): AgentErrorBench is dominated by planning failures (34.5%) while, under the consensus coding, Who&When's largest classes are malformed tool calls (20.1%) and fabricated observations (19.6%), although sensitivity analyses show that the tool-formatting plurality depends on how ambiguous "incorrect code" explanations are classified. Failure profiles also differ significantly across agent domains (p=1.17×10⁻²⁷) and across the three evaluated language models (p=0.002, V=0.30). Contrary to our preregistered expectation, coordination failures do not dominate the multi-agent data. We also audit benchmark label quality directly: 13.5% of AgentErrorBench annotations carry an empty failure label, concentrated in one model's trajectories (42.1% of Llama3.3-70B-Turbo vs. 0% of GPT-4o); the released label files contain inconsistent category strings; and the taxonomy's documented size (17 types) does not match its released definition file (18 distinct strings). No released annotation in either dataset mapped to the no recovery class under our coding procedure. Because the two datasets differ simultaneously in tasks, systems, and annotation protocols, these distributional gaps cannot isolate a taxonomy effect; we frame the study as a descriptive audit of harmonized failure annotations, and report the confounds explicitly. Our results bound what portability claims the public evidence can support, and show that label noise in public failure datasets is itself a first-order problem for the field. We close with concrete recommendations for benchmark designers.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-01
DOI
https://doi.org/10.5281/zenodo.23072906
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

How Well Do Agent Failure Taxonomies Travel? A Cross-Dataset Empirical Analysis of Failure Labels in Public Agent Trajectory Datasets

Tejaswini Viswanath, Abhinav Tharamel
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

How Well Do Agent Failure Taxonomies Travel? A Cross-Dataset Empirical Analysis of Failure Labels in Public Agent Trajectory Datasets

Tejaswini Viswanath, Abhinav Tharamel
preprint en

Abstract

Every published taxonomy of AI agent failures is derived from its authors own trajectories, yet the field routinely treats these taxonomies as if they described a universal phenomenon. We audit that assumption empirically. We collect 384 annotated agent failures from two public datasets: AgentErrorBench (200 trajectories from ALFWorld, GAIA and WebShop, labeled with the AgentErrorTaxonomy) and Who&When (184 multi-agent failures with free-text attributions), and map both to a single 13-class unified scheme under a prospective analysis plan. The harmonized label distributions differ dramatically across datasets (χ²=192.53, df=12, p=1.13×10⁻³⁴; Cramer's V=0.708; labeled-only χ²=182.53, df=11, V=0.723): AgentErrorBench is dominated by planning failures (34.5%) while, under the consensus coding, Who&When's largest classes are malformed tool calls (20.1%) and fabricated observations (19.6%), although sensitivity analyses show that the tool-formatting plurality depends on how ambiguous "incorrect code" explanations are classified. Failure profiles also differ significantly across agent domains (p=1.17×10⁻²⁷) and across the three evaluated language models (p=0.002, V=0.30). Contrary to our preregistered expectation, coordination failures do not dominate the multi-agent data. We also audit benchmark label quality directly: 13.5% of AgentErrorBench annotations carry an empty failure label, concentrated in one model's trajectories (42.1% of Llama3.3-70B-Turbo vs. 0% of GPT-4o); the released label files contain inconsistent category strings; and the taxonomy's documented size (17 types) does not match its released definition file (18 distinct strings). No released annotation in either dataset mapped to the no recovery class under our coding procedure. Because the two datasets differ simultaneously in tasks, systems, and annotation protocols, these distributional gaps cannot isolate a taxonomy effect; we frame the study as a descriptive audit of harmonized failure annotations, and report the confounds explicitly. Our results bound what portability claims the public evidence can support, and show that label noise in public failure datasets is itself a first-order problem for the field. We close with concrete recommendations for benchmark designers.

Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.