Stress-Testing Byzantine Robustness in Multi-Agent LLM Systems: Adaptive Adversaries and an Alternative Filtering Anchor
Large language model agents deployed over peer-to-peer networks can improve reliability through consensus, but faulty or adversarial agents can propagate errors through the same channels. Self-Anchored Consensus (SAC; Lee et al., arXiv:2605.09076) addresses this with a decentralized filter-and-refine protocol: each agent scores its neighbors' responses locally, discards the lowest-scoring ones relative to its own self-score, and refines its answer on the remainder. The original evaluation, however, models the adversary as an agent replaying a fixed incorrect response with falsified confidence — an attack the authors identify as a limitation, noting that adversaries crafted to appear correct to the receiver-side scorer are not addressed by the current filtering mechanism. This project examines that gap. We implement stronger adversarial agents using attack templates designed to mimic the structure of honest responses rather than replaying out-of-distribution output, and we evaluate SAC under this threat model. We further replace SAC's self-score filtering anchor with a median-of-neighbors anchor, on the reasoning that anchoring to one's own score limits how much a weak agent can benefit from more capable neighbors. In preliminary experiments using gpt-4o-mini as strong agents and gpt-3.5-turbo as weak agents at n=7 and F=3 across (F+1)-robust topologies (γ-MERG, complete, and Erdos–Rényi random graphs), the modified protocol substantially increases weak-agent accuracy while also improving strong-agent accuracy — a departure from the original protocol, which preserves strong agents rather than lifting them. A matched-baseline comparison against the original self-anchored protocol on the same problems remains to be run. We additionally observe that reported robustness gains in this line of work are sensitive to sample size: the same method and configuration yield markedly different improvements across the two arXiv versions of the original paper, at n=30 and n=100. Current results use 20 mixed-difficulty problems spanning MATH500 Levels 1–5 and are preliminary; evaluation on the full Level 4 subset, with multiple seeds and an ablation isolating the anchor modification, is in progress.
Authors
- Victoria Chen (ORCID: https://orcid.org/0000-0002-2653-8031)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.22849365
- Primary Topic
- Advanced Graph Neural Networks
- Type
- article
- Field-Weighted Citation Impact
- 0.00