Reported, Not Measured: An Empirical Study of Measurement Defects in LLM and Agent Evaluation Tools

Evaluation tools for language-model applications and agent skills are used as gates: a score decides whether a skill ships, which demonstrations a prompt compiler keeps, or whether a model is promoted. This paper is an exploratory study of the software between evidence that has already been produced and the result an evaluation tool reports. Between 14 and 26 September 2026 the author read the scoring and gating code of seven open-source evaluation tools, chosen by the author's judgement rather than sampled (the agent-skills repository's eval harness and a reference gate script, NVIDIA SkillEvaluator, DSPy, LangSmith, DeepEval, Harbor and MLflow), reproduced thirteen defects, twelve of them offline, and filed each upstream with a public reproduction; twelve carried a proposed fix from the author, five are merged, one more is approved, and the rest were open on 30 September 2026. In each case the tool returned a well-formed result that its evidence did not support. The defects fall into six kinds, derived from the cases: absent evidence scored as a measurement, a result bound to the wrong unit, a proxy credited as the act, stale state read as current, a sign error in a threshold, and a check aimed at a target the artefact never claimed. The same taxonomy is applied to the author's own instrument, Driftproof, using only its public findings: a parser that turned a non-numeric judge reply into a score, a badge that verified generation text but not judge text, and six false passes closed in release 0.12.0. Four published measurements serve as illustrations, not as tests: a pass/fail check and a text-injection evaluation that read the same task differently under different treatments; one eval that returned pass eight times and fail twice at one commit; and two release-day comparisons in which most cells did not separate. From the cases the paper proposes nine record fields, each paired with the check that would read it, and maps each defect to the field and check that could have exposed it; the mapping is a design argument the paper has not yet tested. The distinction the paper draws is between a noisy measurement, which statistics can qualify, and a well-formed number that was never a valid measurement, which statistics cannot repair. The evidence is a catalogue of mechanisms found where the author looked; it is a lower bound on what exists, not a prevalence estimate.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23050795
Primary Topic
Artificial Intelligence in Law
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Reported, Not Measured: An Empirical Study of Measurement Defects in LLM and Agent Evaluation Tools

driftproofhq
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Law
preprint

Reported, Not Measured: An Empirical Study of Measurement Defects in LLM and Agent Evaluation Tools

driftproofhq
preprint en

Abstract

Evaluation tools for language-model applications and agent skills are used as gates: a score decides whether a skill ships, which demonstrations a prompt compiler keeps, or whether a model is promoted. This paper is an exploratory study of the software between evidence that has already been produced and the result an evaluation tool reports. Between 14 and 26 September 2026 the author read the scoring and gating code of seven open-source evaluation tools, chosen by the author's judgement rather than sampled (the agent-skills repository's eval harness and a reference gate script, NVIDIA SkillEvaluator, DSPy, LangSmith, DeepEval, Harbor and MLflow), reproduced thirteen defects, twelve of them offline, and filed each upstream with a public reproduction; twelve carried a proposed fix from the author, five are merged, one more is approved, and the rest were open on 30 September 2026. In each case the tool returned a well-formed result that its evidence did not support. The defects fall into six kinds, derived from the cases: absent evidence scored as a measurement, a result bound to the wrong unit, a proxy credited as the act, stale state read as current, a sign error in a threshold, and a check aimed at a target the artefact never claimed. The same taxonomy is applied to the author's own instrument, Driftproof, using only its public findings: a parser that turned a non-numeric judge reply into a score, a badge that verified generation text but not judge text, and six false passes closed in release 0.12.0. Four published measurements serve as illustrations, not as tests: a pass/fail check and a text-injection evaluation that read the same task differently under different treatments; one eval that returned pass eight times and fail twice at one commit; and two release-day comparisons in which most cells did not separate. From the cases the paper proposes nine record fields, each paired with the check that would read it, and maps each defect to the field and check that could have exposed it; the mapping is a design argument the paper has not yet tested. The distinction the paper draws is between a noisy measurement, which statistics can qualify, and a well-formed number that was never a valid measurement, which statistics cannot repair. The evidence is a catalogue of mechanisms found where the author looked; it is a lower bound on what exists, not a prevalence estimate.

Zenodo (CERN European Organization for Nuclear Research)
Quality Education
Artificial Intelligence in Law
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Reported, Not Measured: An Empirical Study of Measurement Defects in LLM and Agent Evaluation Tools — driftproofhq · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS