ProofPathBench: A Benchmark for Verifiability-Aware Planning by Tool-Using Language Agents

ProofPathBench is an implemented research benchmark for whether tool-using language agents select matched plans that produce independent evidence of external-state outcomes. Version 0.0.1 contains 96 synthetic scenarios across eight domains, 384 task-treatment units, 768 deterministic no-fault plan executions, and nine seeded failure classes. The deposit includes a research preprint and an Apache-2.0 benchmark software artifact. It reports no provider-model behavioral results; fixtures and prospective power simulations are not empirical model results. A blinded human construct-validation protocol is included but has not been run. The work has not undergone peer review. Author: Vamshi Krishna Madhavan, Independent Researcher. Correspondence: [email protected]. Source repository: https://github.com/Vamshi0104/proofpathbench. Project website: https://vamshi0104.github.io/proofpathbench/.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23196924
Primary Topic
AI-based Problem Solving and Planning
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

ProofPathBench: A Benchmark for Verifiability-Aware Planning by Tool-Using Language Agents

Vamshi Krishna Madhavan
Zenodo (CERN European Organization for Nuclear Research)
AI-based Problem Solving and Planning
preprint

ProofPathBench: A Benchmark for Verifiability-Aware Planning by Tool-Using Language Agents

Vamshi Krishna Madhavan
preprint en

Abstract

ProofPathBench is an implemented research benchmark for whether tool-using language agents select matched plans that produce independent evidence of external-state outcomes. Version 0.0.1 contains 96 synthetic scenarios across eight domains, 384 task-treatment units, 768 deterministic no-fault plan executions, and nine seeded failure classes. The deposit includes a research preprint and an Apache-2.0 benchmark software artifact. It reports no provider-model behavioral results; fixtures and prospective power simulations are not empirical model results. A blinded human construct-validation protocol is included but has not been run. The work has not undergone peer review. Author: Vamshi Krishna Madhavan, Independent Researcher. Correspondence: [email protected]. Source repository: https://github.com/Vamshi0104/proofpathbench. Project website: https://vamshi0104.github.io/proofpathbench/.

Zenodo (CERN European Organization for Nuclear Research)
AI-based Problem Solving and Planning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.