What Does a Repair Statistic Measure? A Claim-Indexed Audit of Repository-Level LLM Repair

Repository-level repair systems emit many logged events before a verifier pass. A rate over one event is easy to misread when the report omits the event's operational meaning, opportunity budget, support unit, or verifier boundary. We introduce a claim-indexed measurement model that binds each repair statistic to five coordinates: construct, opportunity design, unit of support, endpoint interpretation, and instrumentation provenance. We instantiate the model in a retrospective audit of 2,400 executions: two 30-task DeepSeek strata with disjoint task IDs and a 30-task GLM boundary stratum that reuses the clean task surface. A field-integrity audit identifies 12 raw GLM action positives that conflict with the patch-application evidence required by the governed-action construct. The harmonized table contains no verifier pass outside the governed-action state, while conditional endpoint yield is 14/199 (7.04%) and 14/204 (6.86%) in the two DeepSeek strata. Under eight retained opportunities per fixed cell, governed-action discovery reaches 40.83% and 46.67%; verifier-pass discovery reaches 4.17% and 1.67%. The 28 passing rows reduce to seven positive cells, three tasks, and two task families. Endpoint controls produce the expected result for all six gold runs, six no-op runs, and nine reference-derived mutants. These findings show that repair evidence cannot be reduced to one success scalar. Field semantics, opportunity, support, verifier scope, and instrumentation determine what a reported number measures and which comparisons it can sustain.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-11
DOI
https://doi.org/10.5281/zenodo.22668842
Primary Topic
Scientific Computing and Data Management
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

What Does a Repair Statistic Measure? A Claim-Indexed Audit of Repository-Level LLM Repair

Yuda Bi, Zhida Qin, Jiangwei Xue
Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management
preprint

What Does a Repair Statistic Measure? A Claim-Indexed Audit of Repository-Level LLM Repair

Yuda Bi, Zhida Qin, Jiangwei Xue
preprint en

Abstract

Repository-level repair systems emit many logged events before a verifier pass. A rate over one event is easy to misread when the report omits the event's operational meaning, opportunity budget, support unit, or verifier boundary. We introduce a claim-indexed measurement model that binds each repair statistic to five coordinates: construct, opportunity design, unit of support, endpoint interpretation, and instrumentation provenance. We instantiate the model in a retrospective audit of 2,400 executions: two 30-task DeepSeek strata with disjoint task IDs and a 30-task GLM boundary stratum that reuses the clean task surface. A field-integrity audit identifies 12 raw GLM action positives that conflict with the patch-application evidence required by the governed-action construct. The harmonized table contains no verifier pass outside the governed-action state, while conditional endpoint yield is 14/199 (7.04%) and 14/204 (6.86%) in the two DeepSeek strata. Under eight retained opportunities per fixed cell, governed-action discovery reaches 40.83% and 46.67%; verifier-pass discovery reaches 4.17% and 1.67%. The 28 passing rows reduce to seven positive cells, three tasks, and two task families. Endpoint controls produce the expected result for all six gold runs, six no-op runs, and nine reference-derived mutants. These findings show that repair evidence cannot be reduced to one success scalar. Field semantics, opportunity, support, verifier scope, and instrumentation determine what a reported number measures and which comparisons it can sustain.

Zenodo (CERN European Organization for Nuclear Research)
Center for Translational Research in Neuroimaging and Data Science (US)
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

What Does a Repair Statistic Measure? A Claim-Indexed Audit of Repository-Level LLM Repair — Yuda Bi, Zhida Qin, et al. · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS