Where the Swarm Pays
A ledger of what multiplying LLM agents buys on 240 hard problems across seven reasoning models from five families, revised three times: after a label audit (on the problems where all seven models were scored wrong, the gold label, the question or the grader was at fault on 45 of 48) and after a truncation audit (an answer cut off by a token cap had been counted as a wrong vote). Plurality does not beat the best member; debate's gain is mostly a second chance for agents that ran out of tokens and shrinks from 4.9 to 2.2 points under a different tie rule; the unanimously wrong items it ends on all look like label faults; roles add no measurable independence; synthesis recovers nothing. The comparator result, five judges picking among candidates, is within a tie-break rule of zero. Version 4 of the note; earlier versions and a comparison: https://alatip.github.io/where-the-swarm-pays/versions/
Authors
- Aziz Latipov (ORCID: https://orcid.org/0009-0003-8938-8223)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23229042
- Primary Topic
- Topic Modeling
- Type
- preprint