Performance Assurance in Pull-Request CI: An Operational Method and Evaluation Framework
Software can remain functionally correct while becoming slower, more memory-intensive, or more expensive to operate. This paper considers the narrow problem of evaluating an individual change early enough to inform a pull-request decision. It specializes established Performance Assurance practice by connecting versioned workloads and data, controlled measurement, reference selection, comparability, effect estimation, minimum effects of interest (MEIs), and CI policy. It documents Perfgate 0.1.0, a Ruby gem that compares independent samples with a historical stored baseline, uses deterministic bootstrap intervals for differences in medians, and separates evidence from advisory-by-default policy. A frozen case study completed 126 comparisons for one Rails application on GitHub-hosted workers. Among 30 A/A comparisons, no regression FAIL occurred. At metric level, 598 of 600 decisions were PASS and two were INCONCLUSIVE. Cross-worker compatibility reservations caused 28 otherwise passing comparisons to be reported as overall WARN. The 95% Wilson upper bound for the false-FAIL proportion was 11.4%. All 96 comparisons containing one of eight fixed source regressions produced FAIL evidence, but their unknown, nonrandomized effect sizes make this a detection result rather than a power estimate at an MEI. Two of 63 preassigned rerun pairs reversed between WARN and INCONCLUSIVE. The study did not evaluate same-worker blocks, paired estimation, application-level interval coverage, or direct merge blocking.
Authors
- Oleksandr Potrakhov
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-20
- DOI
- https://doi.org/10.5281/zenodo.22182335
- Primary Topic
- Software System Performance and Reliability
- Type
- article
- Field-Weighted Citation Impact
- 0.00