Checking the Check: Controls, Corpus Alignment, and Scale Invariance in Evaluation Software
We present a three-check protocol for evaluation software: construct discrimination, observation coverage and decision invariance. We apply it to a fixed-output audit and a separate R² case study. The audit uses prefix slices of 270 QA items and retains 273 output-component-by-dataset records, stratified into native-default, custom and off-label configurations; headline item results use only the 220 answerable items. A record is informative when gold and shifted-gold medians satisfy g+ − g− ≥ 0.20 on a normalized higher-is-better scale; an inert output flags when its median is ≥ g+ − 0.10. Four of 140 informative item records (131 with all six inert arms observed; nine lacking empty-output scores) flagged copied context (2 of 91 native-default, 2 of 29 custom, 0 of 20 off-label): two native containment rules and two custom regex configurations. These are construct matches, not demonstrated specification violations. In a three-item lighteval 0.13.0 diagnostic, changing later hypotheses left native chrF/TER unchanged, while SacreBLEU-backed chrF/TER under a correctly oriented reference layout changed. For two fixed predictors with positive mathematical R², TorchMetrics 1.9.0's absolute sum clamp reversed their order at common scale s = 0.01; under the exact branch model the reversal set is exactly 1/16200 < q ≤ 1/9800, where q = s² and s scales targets and predictions together. Mathematical R² remains invariant; a single output under this branch can tie but cannot strictly reverse an ordering. These distinct observations remain unpooled. The results support release-specific implementation checks and careful separation of score agreement, observation coverage and decision validity. They establish neither framework-wide defect rates nor discovery priority.
Authors
- Jared Condon
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-26
- DOI
- https://doi.org/10.5281/zenodo.22973131
- Primary Topic
- Topic Modeling
- Type
- preprint