What a Pooled Success Total Supports for pass@k: Identification, Unbiased Estimation, and a Jackknife Below the Run Threshold

Procedures whose attempts share state, such as search trees, shared-prefix sampling and self-debugging loops, can be scored by a pass@k computed from the total number of successes over all their attempts. We ask what that pooled total supports. In the model where each of R runs draws a latent success rate and makes n conditionally independent attempts, no function of the pooled total is unbiased for pass@k when R, n and k are all at least 2, whether pass@k is read as the procedure’s chance of success within k attempts or as the success chance of k independent attempts. The distribution of the total still determines the independent-attempt target for every k and the procedure target for k ≤ n; we give the complete classification of when an unbiased function of the total exists. Per-run counts, by contrast, are known to admit unbiased estimators of the procedure target for k ≤ n and of the independent-attempt target for k ≤ R. Below that run threshold, at R = k − 1 runs, the delete-one jackknife over the runs’ first attempts has exact worst-case bias 3.86 percentage points at k = 4 and 2.91 at k = 16, far above the minimax bias 2 · 4^−k, and the same values are lower bounds on the worst case of the jackknife over the runs’ success fractions; yet in ten of eleven recorded benchmark cells from eight designs the jackknife’s mean error against held-out runs is within 0.44 points, and −2.02 in the eleventh. An evaluation that logs only the pooled total, once R, n, k ≥ 2, gives up unbiased estimation that per-run counts would allow.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23154869
Primary Topic
Machine Learning and Data Classification
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

What a Pooled Success Total Supports for pass@k: Identification, Unbiased Estimation, and a Jackknife Below the Run Threshold

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Machine Learning and Data Classification
preprint

What a Pooled Success Total Supports for pass@k: Identification, Unbiased Estimation, and a Jackknife Below the Run Threshold

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

Procedures whose attempts share state, such as search trees, shared-prefix sampling and self-debugging loops, can be scored by a pass@k computed from the total number of successes over all their attempts. We ask what that pooled total supports. In the model where each of R runs draws a latent success rate and makes n conditionally independent attempts, no function of the pooled total is unbiased for pass@k when R, n and k are all at least 2, whether pass@k is read as the procedure’s chance of success within k attempts or as the success chance of k independent attempts. The distribution of the total still determines the independent-attempt target for every k and the procedure target for k ≤ n; we give the complete classification of when an unbiased function of the total exists. Per-run counts, by contrast, are known to admit unbiased estimators of the procedure target for k ≤ n and of the independent-attempt target for k ≤ R. Below that run threshold, at R = k − 1 runs, the delete-one jackknife over the runs’ first attempts has exact worst-case bias 3.86 percentage points at k = 4 and 2.91 at k = 16, far above the minimax bias 2 · 4^−k, and the same values are lower bounds on the worst case of the jackknife over the runs’ success fractions; yet in ten of eleven recorded benchmark cells from eight designs the jackknife’s mean error against held-out runs is within 0.44 points, and −2.02 in the eleventh. An evaluation that logs only the pooled total, once R, n, k ≥ 2, gives up unbiased estimation that per-run counts would allow.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW)
Machine Learning and Data Classification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

What a Pooled Success Total Supports for pass@k: Identification, Unbiased Estimation, and a Jackknife Below the Run Threshold — Ya-Fen Yeh, Guan-Yuan Chen · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS