Finite-sample decision risk in unseeded quantum compiler benchmarking
Draft preprint. Not submitted to a venue. Quantum compiler benchmarks are used to accept or reject changes to production transpilers, but the compilers they measure are stochastic. At the pinned revision of IBM's Benchpress benchmark suite, its Qiskit gym compiles without passing seed_transpiler, so every gate count it reports is one draw from an unmeasured distribution — while the BQSKit gym in the same repository does seed its compiler. We quantify the consequence as finite-sample decision risk: the probability that a regression verdict computed from k runs per version disagrees with the verdict implied by the long-run mean of the same measurement. On bv_n140 mapped to a heavy-hex lattice, measured over 400 seeds per version across 21 OS processes, the change from Qiskit 1.4.3 to 2.0.0 has a long-run mean of +5.37% (95% CI +4.27% to +6.50%). The suite's own protocol — three unseeded runs per version — reports it as a ≥+10% regression 24.4% of the time (95% CI 19.5–31.1%). A disjoint 200-seed sample across ten fresh processes reproduces this at 22.6%. Twenty runs per version, roughly 40 hours of compute, still leaves 3.7%. A pre-registered replication across 39 circuits is reported in full: the analysis code was committed before the data existed, and the commit ordering is included in this archive as COMMIT_TIMESTAMPS.txt. The pre-registered primary endpoint is 12 of 26 eligible circuits (46.2%, Wilson 95% CI 28.8–64.5%), reported alongside a cluster-robust interval of [17.4%, 81.0%] that reflects algorithm-family clustering. What this archive contains: the paper, the complete analysis and measurement code, and all 41,790 raw per-seed measurements. verify.py re-runs the toolchain pin check, test suite, numeric inventory, replication artifact and a proof of a withdrawn analytical claim in under a minute. Four claims from earlier phases of this work were withdrawn on the record after adversarial review, and the withdrawals remain in the repository with the evidence that defeated them. Passing seed_transpiler removes false positives on the circuits tested but was worse on three of four circuits exhibiting false negatives, and is not presented as a general remedy.
Authors
- Panagiotis Gkilis
Institutions
- Bevital (Norway) (NO)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-04
- DOI
- https://doi.org/10.5281/zenodo.22310060
- Primary Topic
- Quantum Computing Algorithms and Architecture
- Type
- preprint