Checking the Check: Controls, Corpus Alignment, and Scale Invariance in Evaluation Software

Evaluation frameworks turn model outputs into the scores that benchmarks, leaderboards and deployment decisions rely on, and the scoring code itself also requires direct checks. We test four evaluation frameworks (DeepEval, TruLens, lighteval and HELM) and one metrics library (TorchMetrics) with three checks: does a score separate a correct answer from a wrong one (construct discrimination), does every supplied observation reach the score (observation coverage), and does a transformation that should not change a ranking leave it unchanged (decision invariance)? Six fixed inert outputs (abstaining, empty, copying the question, copying the context, a constant sentence and refusing) and three controls produce 273 output-component-by-dataset records on 270 QA items. Three findings follow. First, on single-reference corpora, the corpus chrF, chrF++ and TER wrappers in lighteval 0.13.0 pass references to SacreBLEU in an orientation that scores only the first hypothesis, so correct and shifted wrong outputs both received chrF 100 and TER 0, and a three-item diagnostic against a correctly oriented SacreBLEU call confirms the lost observations. Second, TorchMetrics 1.9.0 applies an absolute tolerance to unnormalized sums in R², which reverses the ranking of two fixed predictors with positive R² when targets and predictions are rescaled together by s = 0.01; under the exact branch model, the reversal set is 1/16200 < s² ≤ 1/9800, while mathematical R² and scikit-learn 1.6.1 keep the order. Third, simple inert outputs rarely defeat the deterministic metrics tested: of 140 records whose controls discriminated (131 with all six inert arms observed), 4 flagged, all from copying the context into containment-based scorers, two native and two custom. These are construct matches, not demonstrated specification violations. We give a five-step checklist for maintainers and release all inputs, outputs, pinned source locations and aggregation code. The observations are release-specific; they establish neither framework-wide defect rates nor discovery priority.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-27
DOI
https://doi.org/10.5281/zenodo.22973130
Primary Topic
Topic Modeling
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Checking the Check: Controls, Corpus Alignment, and Scale Invariance in Evaluation Software

Jared Condon
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
preprint

Checking the Check: Controls, Corpus Alignment, and Scale Invariance in Evaluation Software

Jared Condon
preprint en

Abstract

Evaluation frameworks turn model outputs into the scores that benchmarks, leaderboards and deployment decisions rely on, and the scoring code itself also requires direct checks. We test four evaluation frameworks (DeepEval, TruLens, lighteval and HELM) and one metrics library (TorchMetrics) with three checks: does a score separate a correct answer from a wrong one (construct discrimination), does every supplied observation reach the score (observation coverage), and does a transformation that should not change a ranking leave it unchanged (decision invariance)? Six fixed inert outputs (abstaining, empty, copying the question, copying the context, a constant sentence and refusing) and three controls produce 273 output-component-by-dataset records on 270 QA items. Three findings follow. First, on single-reference corpora, the corpus chrF, chrF++ and TER wrappers in lighteval 0.13.0 pass references to SacreBLEU in an orientation that scores only the first hypothesis, so correct and shifted wrong outputs both received chrF 100 and TER 0, and a three-item diagnostic against a correctly oriented SacreBLEU call confirms the lost observations. Second, TorchMetrics 1.9.0 applies an absolute tolerance to unnormalized sums in R², which reverses the ranking of two fixed predictors with positive R² when targets and predictions are rescaled together by s = 0.01; under the exact branch model, the reversal set is 1/16200 < s² ≤ 1/9800, while mathematical R² and scikit-learn 1.6.1 keep the order. Third, simple inert outputs rarely defeat the deterministic metrics tested: of 140 records whose controls discriminated (131 with all six inert arms observed), 4 flagged, all from copying the context into containment-based scorers, two native and two custom. These are construct matches, not demonstrated specification violations. We give a five-step checklist for maintainers and release all inputs, outputs, pinned source locations and aggregation code. The observations are release-specific; they establish neither framework-wide defect rates nor discovery priority.

Zenodo (CERN European Organization for Nuclear Research)
Reduced inequalities
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Checking the Check: Controls, Corpus Alignment, and Scale Invariance in Evaluation Software — Jared Condon · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS