Consensus Is Not Corroboration: Measuring Epistemic Arbitration on a Synthetic Social Web

Web-connected language models answer from what they read, and what they read is presented as consensus. Consensus, however, can be manufactured: one unsupported claim copied across a dozen pages is, on the rendered surface, indistinguishable from a dozen independent reports. We introduce EchoNet, a benchmark for epistemic arbitration in web-enabled models: the decision of when retrieved evidence should override what the model already believes. Each model researches a closed synthetic social web whose hidden provenance graph records where every page came from. One hundred claims are instantiated under six matched counterfactual conditions (poisoned pages, manipulated search rankings, manufactured consensus, legitimate updates, and a false visible majority) that vary only the stance and provenance topology of the pages. Surface statistics are held fixed, and a deterministic scorer grades belief change against ground truth. Pilots across eleven frontier model configurations give four results. First, GLM 5.2 leads the raw EAS ranking at 1.000, ahead of Gemini 3.7 Flash (0.984) and Qwen3.7 Max (0.951). On five of the six conditions, accuracy spans 0.71 to 1.00 with a spread of at most 0.29, while the false-majority condition spreads 0.42, from 0.50 to 0.92, and accounts for most of the observed separation. Second, paired bootstrap on shared episodes places Gemini 3.7 Flash, Gemini 3.5 Flash-Lite, and Muse Spark 1.2 in one statistical cluster. Third, resistance and revision decouple. Four models never abandon a correct prior, yet the field differs five-fold in how often a model disowns the primary source after finding it (0.058 to 0.306). Fourth, inference cost is weakly associated with arbitration quality at list pricing: the cheapest configuration, statistically tied with the leader, costs far less than the most expensive configuration while gaining only 0.046 of EAS over it.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.22875451
Primary Topic
Scientific Computing and Data Management
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Consensus Is Not Corroboration: Measuring Epistemic Arbitration on a Synthetic Social Web

Divyansh Agrawal
Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management
preprint

Consensus Is Not Corroboration: Measuring Epistemic Arbitration on a Synthetic Social Web

Divyansh Agrawal
preprint en

Abstract

Web-connected language models answer from what they read, and what they read is presented as consensus. Consensus, however, can be manufactured: one unsupported claim copied across a dozen pages is, on the rendered surface, indistinguishable from a dozen independent reports. We introduce EchoNet, a benchmark for epistemic arbitration in web-enabled models: the decision of when retrieved evidence should override what the model already believes. Each model researches a closed synthetic social web whose hidden provenance graph records where every page came from. One hundred claims are instantiated under six matched counterfactual conditions (poisoned pages, manipulated search rankings, manufactured consensus, legitimate updates, and a false visible majority) that vary only the stance and provenance topology of the pages. Surface statistics are held fixed, and a deterministic scorer grades belief change against ground truth. Pilots across eleven frontier model configurations give four results. First, GLM 5.2 leads the raw EAS ranking at 1.000, ahead of Gemini 3.7 Flash (0.984) and Qwen3.7 Max (0.951). On five of the six conditions, accuracy spans 0.71 to 1.00 with a spread of at most 0.29, while the false-majority condition spreads 0.42, from 0.50 to 0.92, and accounts for most of the observed separation. Second, paired bootstrap on shared episodes places Gemini 3.7 Flash, Gemini 3.5 Flash-Lite, and Muse Spark 1.2 in one statistical cluster. Third, resistance and revision decouple. Four models never abandon a correct prior, yet the field differs five-fold in how often a model disowns the primary source after finding it (0.058 to 0.306). Fourth, inference cost is weakly associated with arbitration quality at list pricing: the cheapest configuration, statistically tied with the leader, costs far less than the most expensive configuration while gaining only 0.046 of EAS over it.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.