Repair Evidence Resolution: Separating Prediction, Structural Resolution, and Transfer in Automated Patch Correctness Assessment

Automated patch correctness assessment (APCA) is usually evaluated by predictive performance: accuracy, precision, recall, F1, ranking quality, or related measures. These metrics answer whether an assessor tends to make the right decision, but they do not determine whether the evidence visible to the assessor is structurally capable of separating repairs that require opposite correctness decisions. We introduce Repair Evidence Resolution (RER), a claim-relative framework for studying that prior question. RER models assessor-visible evidence as a partition of a declared repair population and asks whether every evidence cell is homogeneous with respect to a repair claim. It yields repair-claim resolution (RCR), concrete claim-conflict witnesses (CCWs), resolution deficit (RD), claim-resolution profiles, and claim-boundary completeness (CBC) for transferring resolution conclusions from evaluated repair subsets to larger target populations. We prove basic separation results and explicitly reduce exact-identification special cases to established teaching/certificate-complexity machinery. We then implement a deterministic audit pipeline and evaluate a publisher-verified ASE 2020 APCA artifact. Across 1,762 generator-aware records, one released evidence channel produces nine mixed correctness cells covering 43 patches (RD=2.44%), whereas a second channel and the combined evidence are exact-cell claim-resolving on the observed corpus. We also derive the exact probability that restricted evaluation subsets falsely appear resolved: for a sample of 100 records it is 91.29%. The results motivate a three-layer evaluation view: prediction, resolution, and transfer.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-25
DOI
https://doi.org/10.5281/zenodo.22961348
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Repair Evidence Resolution: Separating Prediction, Structural Resolution, and Transfer in Automated Patch Correctness Assessment

Md. Amir Khusru Akhtar
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

Repair Evidence Resolution: Separating Prediction, Structural Resolution, and Transfer in Automated Patch Correctness Assessment

Md. Amir Khusru Akhtar
preprint en

Abstract

Automated patch correctness assessment (APCA) is usually evaluated by predictive performance: accuracy, precision, recall, F1, ranking quality, or related measures. These metrics answer whether an assessor tends to make the right decision, but they do not determine whether the evidence visible to the assessor is structurally capable of separating repairs that require opposite correctness decisions. We introduce Repair Evidence Resolution (RER), a claim-relative framework for studying that prior question. RER models assessor-visible evidence as a partition of a declared repair population and asks whether every evidence cell is homogeneous with respect to a repair claim. It yields repair-claim resolution (RCR), concrete claim-conflict witnesses (CCWs), resolution deficit (RD), claim-resolution profiles, and claim-boundary completeness (CBC) for transferring resolution conclusions from evaluated repair subsets to larger target populations. We prove basic separation results and explicitly reduce exact-identification special cases to established teaching/certificate-complexity machinery. We then implement a deterministic audit pipeline and evaluate a publisher-verified ASE 2020 APCA artifact. Across 1,762 generator-aware records, one released evidence channel produces nine mixed correctness cells covering 43 patches (RD=2.44%), whereas a second channel and the combined evidence are exact-cell claim-resolving on the observed corpus. We also derive the exact probability that restricted evaluation subsets falsely appear resolved: for a sample of 100 records it is 91.29%. The results motivate a three-layer evaluation view: prediction, resolution, and transfer.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.