Verification Theater: Measured Failure Modes of AI Verifier Panels

Accepted as a poster at the NeurIPS 2026 workshop Who Verifies the Agents? Toward Reliable Agent Development (Sydney, December 2026). The workshop is non-archival; this deposit is the citable version.The same text is also at https://github.com/vadimchernets/papers/tree/main/verification-theater. Pre-registration and replication package: https://osf.io/4y9dv. Multi-agent verification is becoming a default trust mechanism: an answer is trusted because independent models agreed on it. Using a pre-registered, blind, temperature-zero collection in which sixteen models (nine frontier vendors, seven local open-weights) answered the same 1,500 factual questions (plus a 750-question multiple-choice arm and an exploratory dialogue arm), we measure when panel verification stops working while continuing to look like it works—a condition we call verification theater. We report four failure regimes. First, a difficulty inversion: the certification lift of independent agreement—×2.7–4.6 on the easier three quintiles of questions—falls to 0.82× on the second-hardest quintile and to zero on the hardest, where replicated answers were correct in 0 of 342 cases (≈13 expected under independence). A leave-panel-out re-stratification (difficulty ranked by the 13 non-panel models) removes the pooled inversion (78/354 correct, lift 2.36×) while the zero persists for the cross-bloc panel (0/116); we report both codings together throughout: on the questions the deployed population finds hardest, agreement functions as error certification. Second, constrained answer spaces: under multiple choice (local tier), panels unanimously certified a wrong answer on 4.0–7.9% of questions, versus 1.5–4.6% on the same tier's open-form answers and 0.3–4.9% frontier open-form. Third, the labels used as proxies for independence fail empirically: geopolitical bloc does not predict error correlation (mean intra-bloc φ 0.285 vs. cross-bloc 0.274, difference +0.012), and a reasoning model distilled from the same base as its panel-mate was among the least correlated pairs measured (φ = 0.135; 14th of the 15 US/CN-bloc local pairs, 20th of all 21 local pairs). Fourth, conformity: a verifier that had disagreed with an answer blind endorsed the same answer in 38.5% of cases when shown it as a colleague's; a reverse-prompt control shows these flips track the presence of a confident proposal, not its truth (38.9% toward wrong vs. 34.7% toward correct; McNemar exact p = .012). We argue these regimes share an economic root and derive a measurement discipline that separates verification from its theater. Companion paper: Weakly Discriminative, Not Repaired by Profile Injection, and Manipulable: Self-Confidence in LLM Evaluation (NeurIPS 2026 workshop TAE), 10.5281/zenodo.23197651.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23197641
Citations
1
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Verification Theater: Measured Failure Modes of AI Verifier Panels

Vadym Chernets
1 citations
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

Verification Theater: Measured Failure Modes of AI Verifier Panels

Vadym Chernets
preprint en
1 citations

Abstract

Accepted as a poster at the NeurIPS 2026 workshop Who Verifies the Agents? Toward Reliable Agent Development (Sydney, December 2026). The workshop is non-archival; this deposit is the citable version.The same text is also at https://github.com/vadimchernets/papers/tree/main/verification-theater. Pre-registration and replication package: https://osf.io/4y9dv. Multi-agent verification is becoming a default trust mechanism: an answer is trusted because independent models agreed on it. Using a pre-registered, blind, temperature-zero collection in which sixteen models (nine frontier vendors, seven local open-weights) answered the same 1,500 factual questions (plus a 750-question multiple-choice arm and an exploratory dialogue arm), we measure when panel verification stops working while continuing to look like it works—a condition we call verification theater. We report four failure regimes. First, a difficulty inversion: the certification lift of independent agreement—×2.7–4.6 on the easier three quintiles of questions—falls to 0.82× on the second-hardest quintile and to zero on the hardest, where replicated answers were correct in 0 of 342 cases (≈13 expected under independence). A leave-panel-out re-stratification (difficulty ranked by the 13 non-panel models) removes the pooled inversion (78/354 correct, lift 2.36×) while the zero persists for the cross-bloc panel (0/116); we report both codings together throughout: on the questions the deployed population finds hardest, agreement functions as error certification. Second, constrained answer spaces: under multiple choice (local tier), panels unanimously certified a wrong answer on 4.0–7.9% of questions, versus 1.5–4.6% on the same tier's open-form answers and 0.3–4.9% frontier open-form. Third, the labels used as proxies for independence fail empirically: geopolitical bloc does not predict error correlation (mean intra-bloc φ 0.285 vs. cross-bloc 0.274, difference +0.012), and a reasoning model distilled from the same base as its panel-mate was among the least correlated pairs measured (φ = 0.135; 14th of the 15 US/CN-bloc local pairs, 20th of all 21 local pairs). Fourth, conformity: a verifier that had disagreed with an answer blind endorsed the same answer in 38.5% of cases when shown it as a colleague's; a reverse-prompt control shows these flips track the presence of a confident proposal, not its truth (38.9% toward wrong vs. 34.7% toward correct; McNemar exact p = .012). We argue these regimes share an economic root and derive a measurement discipline that separates verification from its theater. Companion paper: Weakly Discriminative, Not Repaired by Profile Injection, and Manipulable: Self-Confidence in LLM Evaluation (NeurIPS 2026 workshop TAE), 10.5281/zenodo.23197651.

Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.