A Hard Verbatim-Evidence Gate as the Precondition for Multi-Agent Claim Verification: Design, Deployment, and Empirical Boundary Conditions

Multi-agent large language model (LLM) pipelines verify claims by having model instances critique one another, but agreement reached through opinion carries no guarantee that any evidence supports it, and prompt-enforced grounding reduces unsupported assertion without preventing it. This study evaluates a structural alternative: a verdict may be Supported or Contradicted only if it quotes a span that pipeline code finds verbatim in the retrieved evidence; otherwise it is rewritten to Unverifiable. The gate was built into Aletheia, a deployed LangGraph verification service, and compared with a single-LLM baseline and an otherwise identical ungrounded ablation on SciFact and FEVER, using one frozen corpus, seeded stratified samples of 100 claims, and paired significance tests. On the deployed model the grounded arm ties the baseline on accuracy (79.0 versus 79.0 percent) and leads on catch rate (96.6 versus 93.1 percent) and false agreement (6.1 versus 10.5 percent), but neither interval excludes zero. On an 8-billion-parameter model the catch-rate gain is significant: 82.8 versus 60.3 percent, +22.4 points, 95 percent confidence interval [+12.1, +33.3]. On FEVER's paraphrased claims the ungrounded ablation is significantly more accurate (85.0 versus 77.0 percent). Re-checked from traces, 51 of 51 asserted FEVER verdicts are verbatim-backed, yet five are wrong: the gate guarantees a faithful quotation, not a correct inference. Structural grounding buys accuracy for weak models and auditability for strong ones — a boundary condition, not a leaderboard win.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-26
DOI
https://doi.org/10.5281/zenodo.22978022
Primary Topic
Topic Modeling
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

A Hard Verbatim-Evidence Gate as the Precondition for Multi-Agent Claim Verification: Design, Deployment, and Empirical Boundary Conditions

Jay Gautam, Rahul Kumar
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
preprint

A Hard Verbatim-Evidence Gate as the Precondition for Multi-Agent Claim Verification: Design, Deployment, and Empirical Boundary Conditions

Jay Gautam, Rahul Kumar
preprint en

Abstract

Multi-agent large language model (LLM) pipelines verify claims by having model instances critique one another, but agreement reached through opinion carries no guarantee that any evidence supports it, and prompt-enforced grounding reduces unsupported assertion without preventing it. This study evaluates a structural alternative: a verdict may be Supported or Contradicted only if it quotes a span that pipeline code finds verbatim in the retrieved evidence; otherwise it is rewritten to Unverifiable. The gate was built into Aletheia, a deployed LangGraph verification service, and compared with a single-LLM baseline and an otherwise identical ungrounded ablation on SciFact and FEVER, using one frozen corpus, seeded stratified samples of 100 claims, and paired significance tests. On the deployed model the grounded arm ties the baseline on accuracy (79.0 versus 79.0 percent) and leads on catch rate (96.6 versus 93.1 percent) and false agreement (6.1 versus 10.5 percent), but neither interval excludes zero. On an 8-billion-parameter model the catch-rate gain is significant: 82.8 versus 60.3 percent, +22.4 points, 95 percent confidence interval [+12.1, +33.3]. On FEVER's paraphrased claims the ungrounded ablation is significantly more accurate (85.0 versus 77.0 percent). Re-checked from traces, 51 of 51 asserted FEVER verdicts are verbatim-backed, yet five are wrong: the gate guarantees a faithful quotation, not a correct inference. Structural grounding buys accuracy for weak models and auditability for strong ones — a boundary condition, not a leaderboard win.

Zenodo (CERN European Organization for Nuclear Research)
Soka University of America (US)
Peace, Justice and strong institutions
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A Hard Verbatim-Evidence Gate as the Precondition for Multi-Agent Claim Verification: Design, Deployment, and Empirical Boundary Conditions — Jay Gautam, Rahul Kumar · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS