When Too Many Cooks Spoil the Broth: Three Failure Modes of Multi-Agent LLM Reliability

Multi-agent LLM systems are widely deployed under three implicit assumptions: that self-confidence calibration transfers to peer evaluation, that errors made by distinct instances are statistically independent, and that any sufficiently calibrated model is an equivalent verifier. We test all three across GPT-4o, Claude Sonnet 4.6, and Llama-3.3-70B. Self-confidence and peer-resistance dissociate: Claude shows the strongest self-confidence signal yet accepts every deliberately-wrong peer output, the opposite of what unidimensional calibration predicts. Four independent GPT-4o instances agree on the same wrong answer \\(56\\%\\) of the time on obscure factual queries versus \\(0.4\\%\\) predicted under independence—a \\(140\\times\\) violation. A pre-registered cross-family replication ( \\(N=200\\) , three architecturally distinct families) yields \\(55\\%\\) three-way agreement, statistically indistinguishable from the same-family baseline (one-sided binomial \\(p=0.41\\) ): family diversification does not reduce correlated hallucination. Verifier choice produces a 42-percentage-point spread in error propagation. We propose a three-phase framework treating self-confidence assessment, peer-resistance capability, and error correlation as orthogonal axes requiring distinct training signals, and identify query-inherent rarity, not shared training, as the dominant driver of correlated failure.

Authors

Institutions

Publication Details

Journal
ACM AI Letters
Published
2026-09-17
DOI
https://doi.org/10.1145/3847307
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

When Too Many Cooks Spoil the Broth: Three Failure Modes of Multi-Agent LLM Reliability

Aradhana Rai, Vijay K. Madisetti
ACM AI Letters
Adversarial Robustness in Machine Learning
article

When Too Many Cooks Spoil the Broth: Three Failure Modes of Multi-Agent LLM Reliability

Aradhana Rai, Vijay K. Madisetti
article en

Abstract

Multi-agent LLM systems are widely deployed under three implicit assumptions: that self-confidence calibration transfers to peer evaluation, that errors made by distinct instances are statistically independent, and that any sufficiently calibrated model is an equivalent verifier. We test all three across GPT-4o, Claude Sonnet 4.6, and Llama-3.3-70B. Self-confidence and peer-resistance dissociate: Claude shows the strongest self-confidence signal yet accepts every deliberately-wrong peer output, the opposite of what unidimensional calibration predicts. Four independent GPT-4o instances agree on the same wrong answer \(56\%\) of the time on obscure factual queries versus \(0.4\%\) predicted under independence—a \(140\times\) violation. A pre-registered cross-family replication ( \(N=200\) , three architecturally distinct families) yields \(55\%\) three-way agreement, statistically indistinguishable from the same-family baseline (one-sided binomial \(p=0.41\) ): family diversification does not reduce correlated hallucination. Verifier choice produces a 42-percentage-point spread in error propagation. We propose a three-phase framework treating self-confidence assessment, peer-resistance capability, and error correlation as orthogonal axes requiring distinct training signals, and identify query-inherent rarity, not shared training, as the dominant driver of correlated failure.

ACM AI Letters
Georgia Institute of Technology (US)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

When Too Many Cooks Spoil the Broth: Three Failure Modes of Multi-Agent LLM Reliability — Aradhana Rai, Vijay K. Madisetti · ACM AI Letters (2026) | TGRS Research Map | TGRS