What pass@k Measures When Attempts Share State: Count Runs, Not Attempts

Search trees, planners and self-correcting agents make the attempts of one run share a random state, yet their results are often scored with a pass@k estimator built for independent draws, which public harnesses apply to whatever attempts they receive. Prior work separates the procedure's success probability from the independent-attempt one; we ask which of them such a number measures and how to report it validly, taking the complete run as the sampling unit. Whether a run succeeds within k attempts is an unbiased estimate of the procedure's success probability succ@k under any within-run dependence, and succ@k is the only quantity estimable from the first k attempts, with estimates in [0,1], that equals pass@k when attempts are independent. Pooling attempts across runs estimates a third quantity, which we give in closed form and which, for attempts conditionally i.i.d. within a run, lies between succ@k and the independent-attempt target pass@k^iid. Under a calibration condition and over a rich class of dependence structures, pass@k^iid has a uniformly unbiased estimator exactly when at least k complete runs are observed, however much is recorded inside each run, while per-question run-level confidence sequences are valid at any number of runs, for succ@k and, under the same condition, for pass@k^iid. In 8 designs spanning four dependence mechanisms, the pooled report exceeds the run-level estimate of succ@4 in every design, by +0.0205 to +0.1766, and on one design a run-level interval contains a fresh reference 0.952 of the time, against 0.104 for an attempt-level one. The fix is to name the target and count fresh runs.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-25
DOI
https://doi.org/10.5281/zenodo.22959605
Primary Topic
Software Engineering Research
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

What pass@k Measures When Attempts Share State: Count Runs, Not Attempts

Von‐Wun Soo, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Software Engineering Research
preprint

What pass@k Measures When Attempts Share State: Count Runs, Not Attempts

Von‐Wun Soo, Guan-Yuan Chen
preprint en

Abstract

Search trees, planners and self-correcting agents make the attempts of one run share a random state, yet their results are often scored with a pass@k estimator built for independent draws, which public harnesses apply to whatever attempts they receive. Prior work separates the procedure's success probability from the independent-attempt one; we ask which of them such a number measures and how to report it validly, taking the complete run as the sampling unit. Whether a run succeeds within k attempts is an unbiased estimate of the procedure's success probability succ@k under any within-run dependence, and succ@k is the only quantity estimable from the first k attempts, with estimates in [0,1], that equals pass@k when attempts are independent. Pooling attempts across runs estimates a third quantity, which we give in closed form and which, for attempts conditionally i.i.d. within a run, lies between succ@k and the independent-attempt target pass@k^iid. Under a calibration condition and over a rich class of dependence structures, pass@k^iid has a uniformly unbiased estimator exactly when at least k complete runs are observed, however much is recorded inside each run, while per-question run-level confidence sequences are valid at any number of runs, for succ@k and, under the same condition, for pass@k^iid. In 8 designs spanning four dependence mechanisms, the pooled report exceeds the run-level estimate of succ@4 in every design, by +0.0205 to +0.1766, and on one design a run-level interval contains a fresh reference 0.952 of the time, against 0.104 for an attempt-level one. The fix is to name the target and count fresh runs.

Zenodo (CERN European Organization for Nuclear Research)
Chang Gung University (TW), National Tsing Hua University (TW), North Carolina Exploring Cultural Heritage Online (US)
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

What pass@k Measures When Attempts Share State: Count Runs, Not Attempts — Von‐Wun Soo, Guan-Yuan Chen · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS