Finite-sample decision risk in unseeded quantum compiler benchmarking

Draft preprint. Not submitted to a venue. Quantum compiler benchmarks are used to accept or reject changes to production transpilers, but the compilers they measure are stochastic. At the pinned revision of IBM's Benchpress benchmark suite, its Qiskit gym compiles without passing seed_transpiler, so every gate count it reports is one draw from an unmeasured distribution — while the BQSKit gym in the same repository does seed its compiler. We quantify the consequence as finite-sample decision risk: the probability that a regression verdict computed from k runs per version disagrees with the verdict implied by the long-run mean of the same measurement. On bv_n140 mapped to a heavy-hex lattice, measured over 400 seeds per version across 21 OS processes, the change from Qiskit 1.4.3 to 2.0.0 has a long-run mean of +5.37% (95% CI +4.27% to +6.50%). The suite's own protocol — three unseeded runs per version — reports it as a ≥+10% regression 24.4% of the time (95% CI 19.5–31.1%). A disjoint 200-seed sample across ten fresh processes reproduces this at 22.6%. Twenty runs per version, roughly 40 hours of compute, still leaves 3.7%. A pre-registered replication across 39 circuits is reported in full: the analysis code was committed before the data existed, and the commit ordering is included in this archive as COMMIT_TIMESTAMPS.txt. The pre-registered primary endpoint is 12 of 26 eligible circuits (46.2%, Wilson 95% CI 28.8–64.5%), reported alongside a cluster-robust interval of [17.4%, 81.0%] that reflects algorithm-family clustering. What this archive contains: the paper, the complete analysis and measurement code, and all 41,790 raw per-seed measurements. verify.py re-runs the toolchain pin check, test suite, numeric inventory, replication artifact and a proof of a withdrawn analytical claim in under a minute. Four claims from earlier phases of this work were withdrawn on the record after adversarial review, and the withdrawals remain in the repository with the evidence that defeated them. Passing seed_transpiler removes false positives on the circuits tested but was worse on three of four circuits exhibiting false negatives, and is not presented as a general remedy.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-04
DOI
https://doi.org/10.5281/zenodo.22310060
Primary Topic
Quantum Computing Algorithms and Architecture
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Finite-sample decision risk in unseeded quantum compiler benchmarking

Panagiotis Gkilis
Zenodo (CERN European Organization for Nuclear Research)
Quantum Computing Algorithms and Architecture
preprint

Finite-sample decision risk in unseeded quantum compiler benchmarking

Panagiotis Gkilis
preprint en

Abstract

Draft preprint. Not submitted to a venue. Quantum compiler benchmarks are used to accept or reject changes to production transpilers, but the compilers they measure are stochastic. At the pinned revision of IBM's Benchpress benchmark suite, its Qiskit gym compiles without passing seed_transpiler, so every gate count it reports is one draw from an unmeasured distribution — while the BQSKit gym in the same repository does seed its compiler. We quantify the consequence as finite-sample decision risk: the probability that a regression verdict computed from k runs per version disagrees with the verdict implied by the long-run mean of the same measurement. On bv_n140 mapped to a heavy-hex lattice, measured over 400 seeds per version across 21 OS processes, the change from Qiskit 1.4.3 to 2.0.0 has a long-run mean of +5.37% (95% CI +4.27% to +6.50%). The suite's own protocol — three unseeded runs per version — reports it as a ≥+10% regression 24.4% of the time (95% CI 19.5–31.1%). A disjoint 200-seed sample across ten fresh processes reproduces this at 22.6%. Twenty runs per version, roughly 40 hours of compute, still leaves 3.7%. A pre-registered replication across 39 circuits is reported in full: the analysis code was committed before the data existed, and the commit ordering is included in this archive as COMMIT_TIMESTAMPS.txt. The pre-registered primary endpoint is 12 of 26 eligible circuits (46.2%, Wilson 95% CI 28.8–64.5%), reported alongside a cluster-robust interval of [17.4%, 81.0%] that reflects algorithm-family clustering. What this archive contains: the paper, the complete analysis and measurement code, and all 41,790 raw per-seed measurements. verify.py re-runs the toolchain pin check, test suite, numeric inventory, replication artifact and a proof of a withdrawn analytical claim in under a minute. Four claims from earlier phases of this work were withdrawn on the record after adversarial review, and the withdrawals remain in the repository with the evidence that defeated them. Passing seed_transpiler removes false positives on the circuits tested but was worse on three of four circuits exhibiting false negatives, and is not presented as a general remedy.

Zenodo (CERN European Organization for Nuclear Research)
Bevital (Norway) (NO)
Peace, Justice and strong institutions
Quantum Computing Algorithms and Architecture
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.