Checking the Check: Controls, Corpus Alignment, and Scale Invariance in Evaluation Software

We present a three-check protocol for evaluation software: construct discrimination, observation coverage and decision invariance. We apply it to a fixed-output audit and a separate R² case study. The audit uses prefix slices of 270 QA items and retains 273 output-component-by-dataset records, stratified into native-default, custom and off-label configurations; headline item results use only the 220 answerable items. A record is informative when gold and shifted-gold medians satisfy g+ − g− ≥ 0.20 on a normalized higher-is-better scale; an inert output flags when its median is ≥ g+ − 0.10. Four of 140 informative item records (131 with all six inert arms observed; nine lacking empty-output scores) flagged copied context (2 of 91 native-default, 2 of 29 custom, 0 of 20 off-label): two native containment rules and two custom regex configurations. These are construct matches, not demonstrated specification violations. In a three-item lighteval 0.13.0 diagnostic, changing later hypotheses left native chrF/TER unchanged, while SacreBLEU-backed chrF/TER under a correctly oriented reference layout changed. For two fixed predictors with positive mathematical R², TorchMetrics 1.9.0's absolute sum clamp reversed their order at common scale s = 0.01; under the exact branch model the reversal set is exactly 1/16200 < q ≤ 1/9800, where q = s² and s scales targets and predictions together. Mathematical R² remains invariant; a single output under this branch can tie but cannot strictly reverse an ordering. These distinct observations remain unpooled. The results support release-specific implementation checks and careful separation of score agreement, observation coverage and decision validity. They establish neither framework-wide defect rates nor discovery priority.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-26
DOI
https://doi.org/10.5281/zenodo.22973131
Primary Topic
Topic Modeling
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Checking the Check: Controls, Corpus Alignment, and Scale Invariance in Evaluation Software

Jared Condon
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
preprint

Checking the Check: Controls, Corpus Alignment, and Scale Invariance in Evaluation Software

Jared Condon
preprint en

Abstract

We present a three-check protocol for evaluation software: construct discrimination, observation coverage and decision invariance. We apply it to a fixed-output audit and a separate R² case study. The audit uses prefix slices of 270 QA items and retains 273 output-component-by-dataset records, stratified into native-default, custom and off-label configurations; headline item results use only the 220 answerable items. A record is informative when gold and shifted-gold medians satisfy g+ − g− ≥ 0.20 on a normalized higher-is-better scale; an inert output flags when its median is ≥ g+ − 0.10. Four of 140 informative item records (131 with all six inert arms observed; nine lacking empty-output scores) flagged copied context (2 of 91 native-default, 2 of 29 custom, 0 of 20 off-label): two native containment rules and two custom regex configurations. These are construct matches, not demonstrated specification violations. In a three-item lighteval 0.13.0 diagnostic, changing later hypotheses left native chrF/TER unchanged, while SacreBLEU-backed chrF/TER under a correctly oriented reference layout changed. For two fixed predictors with positive mathematical R², TorchMetrics 1.9.0's absolute sum clamp reversed their order at common scale s = 0.01; under the exact branch model the reversal set is exactly 1/16200 < q ≤ 1/9800, where q = s² and s scales targets and predictions together. Mathematical R² remains invariant; a single output under this branch can tie but cannot strictly reverse an ordering. These distinct observations remain unpooled. The results support release-specific implementation checks and careful separation of score agreement, observation coverage and decision validity. They establish neither framework-wide defect rates nor discovery priority.

Zenodo (CERN European Organization for Nuclear Research)
Reduced inequalities, Peace, Justice and strong institutions
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Checking the Check: Controls, Corpus Alignment, and Scale Invariance in Evaluation Software — Jared Condon · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS