Performance Assurance in Pull-Request CI: An Operational Method and Evaluation Framework

Software can remain functionally correct while becoming slower, more memory-intensive, or more expensive to operate. This paper considers the narrow problem of evaluating an individual change early enough to inform a pull-request decision. It specializes established Performance Assurance practice by connecting versioned workloads and data, controlled measurement, reference selection, comparability, effect estimation, minimum effects of interest (MEIs), and CI policy. It documents Perfgate 0.1.0, a Ruby gem that compares independent samples with a historical stored baseline, uses deterministic bootstrap intervals for differences in medians, and separates evidence from advisory-by-default policy. A frozen case study completed 126 comparisons for one Rails application on GitHub-hosted workers. Among 30 A/A comparisons, no regression FAIL occurred. At metric level, 598 of 600 decisions were PASS and two were INCONCLUSIVE. Cross-worker compatibility reservations caused 28 otherwise passing comparisons to be reported as overall WARN. The 95% Wilson upper bound for the false-FAIL proportion was 11.4%. All 96 comparisons containing one of eight fixed source regressions produced FAIL evidence, but their unknown, nonrandomized effect sizes make this a detection result rather than a power estimate at an MEI. Two of 63 preassigned rerun pairs reversed between WARN and INCONCLUSIVE. The study did not evaluate same-worker blocks, paired estimation, application-level interval coverage, or direct merge blocking.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-20
DOI
https://doi.org/10.5281/zenodo.22182335
Primary Topic
Software System Performance and Reliability
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Performance Assurance in Pull-Request CI: An Operational Method and Evaluation Framework

Oleksandr Potrakhov
Zenodo (CERN European Organization for Nuclear Research)
Software System Performance and Reliability
article

Performance Assurance in Pull-Request CI: An Operational Method and Evaluation Framework

Oleksandr Potrakhov
article en

Abstract

Software can remain functionally correct while becoming slower, more memory-intensive, or more expensive to operate. This paper considers the narrow problem of evaluating an individual change early enough to inform a pull-request decision. It specializes established Performance Assurance practice by connecting versioned workloads and data, controlled measurement, reference selection, comparability, effect estimation, minimum effects of interest (MEIs), and CI policy. It documents Perfgate 0.1.0, a Ruby gem that compares independent samples with a historical stored baseline, uses deterministic bootstrap intervals for differences in medians, and separates evidence from advisory-by-default policy. A frozen case study completed 126 comparisons for one Rails application on GitHub-hosted workers. Among 30 A/A comparisons, no regression FAIL occurred. At metric level, 598 of 600 decisions were PASS and two were INCONCLUSIVE. Cross-worker compatibility reservations caused 28 otherwise passing comparisons to be reported as overall WARN. The 95% Wilson upper bound for the false-FAIL proportion was 11.4%. All 96 comparisons containing one of eight fixed source regressions produced FAIL evidence, but their unknown, nonrandomized effect sizes make this a detection result rather than a power estimate at an MEI. Two of 63 preassigned rerun pairs reversed between WARN and INCONCLUSIVE. The study did not evaluate same-worker blocks, paired estimation, application-level interval coverage, or direct merge blocking.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 8%
Software System Performance and Reliability
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Performance Assurance in Pull-Request CI: An Operational Method and Evaluation Framework — Oleksandr Potrakhov · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS