Certifying Data Fitness for Use under Partial Identification
A dataset may satisfy every one of its quality controls and still lead to a decision different from the one the reference state would have produced. The quantity that changes is not a property of the data: it is a loss, relative to a declared use, defined on the coupling between what the system records and what should have been referred to. That coupling, however, is not observed. The central object of this paper is therefore a dissociation: the value of the usage gap may fail to be identified while its adequacy status is. We define the gap, establish that it is in general identified only up to an interval, and show that this interval suffices to decide as soon as it falls on one side of the declared budget — without any need to identify the degradation mechanism. Three observation regimes generate nested identified sets: the recorded state alone yields closed-form bounds, the reference law yields the bounds of a transport problem, and a matched audit yields the value. Verdicts of adequacy and inadequacy are stable along this hierarchy and only indeterminacy is resolved by it, which separates an impossibility of identification from an insufficiency of apparatus: at a fixed observation regime, increasing the sample size reduces estimation uncertainty without reducing identification uncertainty. Two exact decompositions of the gap — by decision classes, then by remediation levers — complete the picture. The paper is theoretical: numerical instances are constructed and declared as such.
Authors
- DIAW (ORCID: https://orcid.org/0009-0008-8463-8525)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22949410
- Primary Topic
- Data Quality and Management
- Type
- preprint