The Perfect Score With a Missing Case: Metric Qualification Does Not Establish Evaluation-Set Conformance
A metric can be correctly computed over the records it receives while those records do not conform to the evaluation set that was declared. This paper reports a preserved evaluator-development lineage completed before target-model release. In V12, the intended nine mechanical and four human-facing measures were implemented, the all-13 self-test passed, and targeted perturbation checks passed, yet direct functional audit showed that one declared required deterministic member could be absent while deterministic_hash_equality still returned 1.0; malformed duplicate or missing grade geometries could also be accepted. The successor V13 placed declared-run identity, exact cardinality, uniqueness, required membership, presentation mapping where blinding applied, and required grade identity/cardinality ahead of metric computation. On a complete 192-run-record functional baseline, the successor gate accepted the control geometry and rejected the tested missing/duplicate attacks. The contribution is not a new checklist of validation primitives: those primitives are substantially occupied by data validation, benchmark auditing, and mature benchmark systems. The bounded result is narrower: under the qualification conditions actually observed in V12, metric qualification was insufficient to establish evaluation-set conformance. V13 is a targeted post-failure repair of the observed attack classes, not evidence of universal sufficiency, prevalence, or model-performance improvement.
Authors
- Logan Davis (ORCID: https://orcid.org/0009-0006-8244-5610)
Institutions
- Logan Regional Hospital (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-06
- DOI
- https://doi.org/10.5281/zenodo.22541028
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- preprint