Stress-Testing Byzantine Robustness in Multi-Agent LLM Systems: Adaptive Adversaries and an Alternative Filtering Anchor

Large language model agents deployed over peer-to-peer networks can improve reliability through consensus, but faulty or adversarial agents can propagate errors through the same channels. Self-Anchored Consensus (SAC; Lee et al., arXiv:2605.09076) addresses this with a decentralized filter-and-refine protocol: each agent scores its neighbors' responses locally, discards the lowest-scoring ones relative to its own self-score, and refines its answer on the remainder. The original evaluation, however, models the adversary as an agent replaying a fixed incorrect response with falsified confidence — an attack the authors identify as a limitation, noting that adversaries crafted to appear correct to the receiver-side scorer are not addressed by the current filtering mechanism. This project examines that gap. We implement stronger adversarial agents using attack templates designed to mimic the structure of honest responses rather than replaying out-of-distribution output, and we evaluate SAC under this threat model. We further replace SAC's self-score filtering anchor with a median-of-neighbors anchor, on the reasoning that anchoring to one's own score limits how much a weak agent can benefit from more capable neighbors. In preliminary experiments using gpt-4o-mini as strong agents and gpt-3.5-turbo as weak agents at n=7 and F=3 across (F+1)-robust topologies (γ-MERG, complete, and Erdos–Rényi random graphs), the modified protocol substantially increases weak-agent accuracy while also improving strong-agent accuracy — a departure from the original protocol, which preserves strong agents rather than lifting them. A matched-baseline comparison against the original self-anchored protocol on the same problems remains to be run. We additionally observe that reported robustness gains in this line of work are sensitive to sample size: the same method and configuration yield markedly different improvements across the two arXiv versions of the original paper, at n=30 and n=100. Current results use 20 mixed-difficulty problems spanning MATH500 Levels 1–5 and are preliminary; evaluation on the full Level 4 subset, with multiple seeds and an ablation isolating the anchor modification, is in progress.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-19
DOI
https://doi.org/10.5281/zenodo.22849365
Primary Topic
Advanced Graph Neural Networks
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Stress-Testing Byzantine Robustness in Multi-Agent LLM Systems: Adaptive Adversaries and an Alternative Filtering Anchor

Victoria Chen
Zenodo (CERN European Organization for Nuclear Research)
Advanced Graph Neural Networks
article

Stress-Testing Byzantine Robustness in Multi-Agent LLM Systems: Adaptive Adversaries and an Alternative Filtering Anchor

Victoria Chen
article en

Abstract

Large language model agents deployed over peer-to-peer networks can improve reliability through consensus, but faulty or adversarial agents can propagate errors through the same channels. Self-Anchored Consensus (SAC; Lee et al., arXiv:2605.09076) addresses this with a decentralized filter-and-refine protocol: each agent scores its neighbors' responses locally, discards the lowest-scoring ones relative to its own self-score, and refines its answer on the remainder. The original evaluation, however, models the adversary as an agent replaying a fixed incorrect response with falsified confidence — an attack the authors identify as a limitation, noting that adversaries crafted to appear correct to the receiver-side scorer are not addressed by the current filtering mechanism. This project examines that gap. We implement stronger adversarial agents using attack templates designed to mimic the structure of honest responses rather than replaying out-of-distribution output, and we evaluate SAC under this threat model. We further replace SAC's self-score filtering anchor with a median-of-neighbors anchor, on the reasoning that anchoring to one's own score limits how much a weak agent can benefit from more capable neighbors. In preliminary experiments using gpt-4o-mini as strong agents and gpt-3.5-turbo as weak agents at n=7 and F=3 across (F+1)-robust topologies (γ-MERG, complete, and Erdos–Rényi random graphs), the modified protocol substantially increases weak-agent accuracy while also improving strong-agent accuracy — a departure from the original protocol, which preserves strong agents rather than lifting them. A matched-baseline comparison against the original self-anchored protocol on the same problems remains to be run. We additionally observe that reported robustness gains in this line of work are sensitive to sample size: the same method and configuration yield markedly different improvements across the two arXiv versions of the original paper, at n=30 and n=100. Current results use 20 mixed-difficulty problems spanning MATH500 Levels 1–5 and are preliminary; evaluation on the full Level 4 subset, with multiple seeds and an ablation isolating the anchor modification, is in progress.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 8%
Advanced Graph Neural Networks
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Stress-Testing Byzantine Robustness in Multi-Agent LLM Systems: Adaptive Adversaries and an Alternative Filtering Anchor — Victoria Chen · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS