What a Data Quality Dashboard Can, and Cannot, Certify
Every indicator of an evaluation instrument can remain strictly unchanged while the decision rule adopted on the basis of that instrument ceases to be optimal. A data quality dashboard is the instance that motivates this work; the framework does not depend on it. The issue is not full-state reconstruction as such. It is whether the instrument observes the information that determines the risk gaps between the decisions actually admissible for a given use. We therefore consider an object whose state is only partially recorded, and an instrument that evaluates it by computing finitely many functions of the recorded part. We formalise such an instrument as a finite-rank linear operator on the space of signed measures. A decision is not a single action applied to an entire population: it is a rule that assigns an action to the recorded information. We accordingly represent a state as a pair consisting of the recorded part and the reference part, and a use as a finite family of admissible rules together with a cost function. The consequences of a rule depend on the coupling between these two coordinates, and not on the law of what is recorded alone. An elementary blindness theorem follows. In the finite-state setting of that theorem, as soon as a cost gap between two admissible rules lies outside the span of the constant function and the checks, there are two regimes the instrument cannot tell apart, between which that gap varies. The instrument’s readings alone can then provide no non-trivial uniform certificate of that variation. A second result separates the existence of an invisible variation from its ability to alter a trade-off. When the optimal rule under the observed state is unique and the decision divergence stays strictly below the margin that separates it from its admissible competitors, that rule remains the unique optimal rule among the admissible ones under the reference state — and, whether or not the certificate is obtained, the regret is bounded by that divergence. The margin depends on the family of admissible rules, and we show why this dependence is not a matter of notational convenience. The relevant quantity here is the divergence generated by the cost gaps between rules. This construction belongs to a lineage of task-oriented divergences; we specifically justify the choice of the class of loss differences by its invariance under the addition of a cost common to all rules and by its direct link with changes of trade-off. We then show that this class pulls back exactly along a pipeline: the divergence measured at the output is exactly the one generated, at the source, by the loss class pulled back along the transformations. Decision quality therefore does not propagate as an intrinsic score from upstream to downstream; the class that defines it is induced by the downstream use and pulled back upstream. The usual Lipschitz bounds then appear as computable relaxations of this exact identity — under regularity conditions that constrain the rules as well, and which we make explicit. We finally state five open problems.
Authors
- Elhadji Thierno DIAW
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22949718
- Primary Topic
- Data Quality and Management
- Type
- preprint