Verification Theater: Measured Failure Modes of AI Verifier Panels
Accepted as a poster at the NeurIPS 2026 workshop Who Verifies the Agents? Toward Reliable Agent Development (Sydney, December 2026). The workshop is non-archival; this deposit is the citable version.The same text is also at https://github.com/vadimchernets/papers/tree/main/verification-theater. Pre-registration and replication package: https://osf.io/4y9dv. Multi-agent verification is becoming a default trust mechanism: an answer is trusted because independent models agreed on it. Using a pre-registered, blind, temperature-zero collection in which sixteen models (nine frontier vendors, seven local open-weights) answered the same 1,500 factual questions (plus a 750-question multiple-choice arm and an exploratory dialogue arm), we measure when panel verification stops working while continuing to look like it works—a condition we call verification theater. We report four failure regimes. First, a difficulty inversion: the certification lift of independent agreement—×2.7–4.6 on the easier three quintiles of questions—falls to 0.82× on the second-hardest quintile and to zero on the hardest, where replicated answers were correct in 0 of 342 cases (≈13 expected under independence). A leave-panel-out re-stratification (difficulty ranked by the 13 non-panel models) removes the pooled inversion (78/354 correct, lift 2.36×) while the zero persists for the cross-bloc panel (0/116); we report both codings together throughout: on the questions the deployed population finds hardest, agreement functions as error certification. Second, constrained answer spaces: under multiple choice (local tier), panels unanimously certified a wrong answer on 4.0–7.9% of questions, versus 1.5–4.6% on the same tier's open-form answers and 0.3–4.9% frontier open-form. Third, the labels used as proxies for independence fail empirically: geopolitical bloc does not predict error correlation (mean intra-bloc φ 0.285 vs. cross-bloc 0.274, difference +0.012), and a reasoning model distilled from the same base as its panel-mate was among the least correlated pairs measured (φ = 0.135; 14th of the 15 US/CN-bloc local pairs, 20th of all 21 local pairs). Fourth, conformity: a verifier that had disagreed with an answer blind endorsed the same answer in 38.5% of cases when shown it as a colleague's; a reverse-prompt control shows these flips track the presence of a confident proposal, not its truth (38.9% toward wrong vs. 34.7% toward correct; McNemar exact p = .012). We argue these regimes share an economic root and derive a measurement discipline that separates verification from its theater. Companion paper: Weakly Discriminative, Not Repaired by Profile Injection, and Manipulable: Self-Confidence in LLM Evaluation (NeurIPS 2026 workshop TAE), 10.5281/zenodo.23197651.
Authors
- Vadym Chernets (ORCID: https://orcid.org/0009-0007-4845-3163)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23197641
- Citations
- 1
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint