High Coverage Does Not Guarantee Regression Detection: A Target-Blind Audit of Representative Task Panels for Versioned AI Coding Agents

Re-evaluating versioned artificial intelligence (AI) coding agents is expensive, motivating small task panels that represent a full benchmark. We audit whether such representativeness predicts regression-screening value (a regression is a task the previous configuration solved and the target failed), with every selector frozen before target outcomes are read. We study Risk-weighted Representative Regression Panels (R3P), which combine a semantic, structural, and behavioral multi-cover core with a stratified probability audit, against nine target-blind baselines on SWE-bench Verified at a 10% task budget. In five declared configuration pairs, R3P attained the highest normalized coverage score (0.840) but estimated the regression rate less accurately than uniform random (mean absolute error (MAE) 0.0639 versus 0.0403) and fragility-stratified sampling (0.0384). The deficit held in all 18 pairs of an exhaustive same-organization cohort and was matched by an equal-sized random core, so under the certainty-core estimator, it reflects budget allocation, not core choice. Under its declared fragility-first seed rule, R3P detected more regressions than random sampling (recall 0.146 versus 0.101), but not through coverage: its fragility-ordered seed held 12 of the 13 regressions its cores caught, and over 500 random seed-visiting orders, median recall was 0.095, below random. High coverage neither guaranteed regression detection nor substituted for probability sampling.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-10-09
DOI
https://doi.org/10.3390/electronics15204597
Primary Topic
Software Engineering Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

High Coverage Does Not Guarantee Regression Detection: A Target-Blind Audit of Representative Task Panels for Versioned AI Coding Agents

Haowei Wang, Qi Cao, Ziran Zhou, Ting Guo
Electronics
Software Engineering Research
article

High Coverage Does Not Guarantee Regression Detection: A Target-Blind Audit of Representative Task Panels for Versioned AI Coding Agents

Haowei Wang, Qi Cao, Ziran Zhou, Ting Guo
article en

Abstract

Re-evaluating versioned artificial intelligence (AI) coding agents is expensive, motivating small task panels that represent a full benchmark. We audit whether such representativeness predicts regression-screening value (a regression is a task the previous configuration solved and the target failed), with every selector frozen before target outcomes are read. We study Risk-weighted Representative Regression Panels (R3P), which combine a semantic, structural, and behavioral multi-cover core with a stratified probability audit, against nine target-blind baselines on SWE-bench Verified at a 10% task budget. In five declared configuration pairs, R3P attained the highest normalized coverage score (0.840) but estimated the regression rate less accurately than uniform random (mean absolute error (MAE) 0.0639 versus 0.0403) and fragility-stratified sampling (0.0384). The deficit held in all 18 pairs of an exhaustive same-organization cohort and was matched by an equal-sized random core, so under the certainty-core estimator, it reflects budget allocation, not core choice. Under its declared fragility-first seed rule, R3P detected more regressions than random sampling (recall 0.146 versus 0.101), but not through coverage: its fragility-ordered seed held 12 of the 13 regressions its cores caught, and over 500 random seed-visiting orders, median recall was 0.095, below random. High coverage neither guaranteed regression detection nor substituted for probability sampling.

ElectronicsVol. 15(20)
National University of Singapore (SG), Hong Kong University of Science and Technology (HK), University of Macau (MO), Ajou University (KR), University of Hong Kong (HK), Anyang University (KR)
Openalex Percentile: Top 6%
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

High Coverage Does Not Guarantee Regression Detection: A Target-Blind Audit of Representative Task Panels for Versioned AI Coding Agents — Haowei Wang, Qi Cao, et al. · Electronics (2026) | TGRS Research Map | TGRS