What a Data Quality Dashboard Can, and Cannot, Certify

Every indicator of an evaluation instrument can remain strictly unchanged while the decision rule adopted on the basis of that instrument ceases to be optimal. A data quality dashboard is the instance that motivates this work; the framework does not depend on it. The issue is not full-state reconstruction as such. It is whether the instrument observes the information that determines the risk gaps between the decisions actually admissible for a given use. We therefore consider an object whose state is only partially recorded, and an instrument that evaluates it by computing finitely many functions of the recorded part. We formalise such an instrument as a finite-rank linear operator on the space of signed measures. A decision is not a single action applied to an entire population: it is a rule that assigns an action to the recorded information. We accordingly represent a state as a pair consisting of the recorded part and the reference part, and a use as a finite family of admissible rules together with a cost function. The consequences of a rule depend on the coupling between these two coordinates, and not on the law of what is recorded alone. An elementary blindness theorem follows. In the finite-state setting of that theorem, as soon as a cost gap between two admissible rules lies outside the span of the constant function and the checks, there are two regimes the instrument cannot tell apart, between which that gap varies. The instrument’s readings alone can then provide no non-trivial uniform certificate of that variation. A second result separates the existence of an invisible variation from its ability to alter a trade-off. When the optimal rule under the observed state is unique and the decision divergence stays strictly below the margin that separates it from its admissible competitors, that rule remains the unique optimal rule among the admissible ones under the reference state — and, whether or not the certificate is obtained, the regret is bounded by that divergence. The margin depends on the family of admissible rules, and we show why this dependence is not a matter of notational convenience. The relevant quantity here is the divergence generated by the cost gaps between rules. This construction belongs to a lineage of task-oriented divergences; we specifically justify the choice of the class of loss differences by its invariance under the addition of a cost common to all rules and by its direct link with changes of trade-off. We then show that this class pulls back exactly along a pipeline: the divergence measured at the output is exactly the one generated, at the source, by the loss class pulled back along the transformations. Decision quality therefore does not propagate as an intrinsic score from upstream to downstream; the class that defines it is induced by the downstream use and pulled back upstream. The usual Lipschitz bounds then appear as computable relaxations of this exact identity — under regularity conditions that constrain the rules as well, and which we make explicit. We finally state five open problems.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-25
DOI
https://doi.org/10.5281/zenodo.22949718
Primary Topic
Data Quality and Management
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

What a Data Quality Dashboard Can, and Cannot, Certify

Elhadji Thierno DIAW
Zenodo (CERN European Organization for Nuclear Research)
Data Quality and Management
preprint

What a Data Quality Dashboard Can, and Cannot, Certify

Elhadji Thierno DIAW
preprint en

Abstract

Every indicator of an evaluation instrument can remain strictly unchanged while the decision rule adopted on the basis of that instrument ceases to be optimal. A data quality dashboard is the instance that motivates this work; the framework does not depend on it. The issue is not full-state reconstruction as such. It is whether the instrument observes the information that determines the risk gaps between the decisions actually admissible for a given use. We therefore consider an object whose state is only partially recorded, and an instrument that evaluates it by computing finitely many functions of the recorded part. We formalise such an instrument as a finite-rank linear operator on the space of signed measures. A decision is not a single action applied to an entire population: it is a rule that assigns an action to the recorded information. We accordingly represent a state as a pair consisting of the recorded part and the reference part, and a use as a finite family of admissible rules together with a cost function. The consequences of a rule depend on the coupling between these two coordinates, and not on the law of what is recorded alone. An elementary blindness theorem follows. In the finite-state setting of that theorem, as soon as a cost gap between two admissible rules lies outside the span of the constant function and the checks, there are two regimes the instrument cannot tell apart, between which that gap varies. The instrument’s readings alone can then provide no non-trivial uniform certificate of that variation. A second result separates the existence of an invisible variation from its ability to alter a trade-off. When the optimal rule under the observed state is unique and the decision divergence stays strictly below the margin that separates it from its admissible competitors, that rule remains the unique optimal rule among the admissible ones under the reference state — and, whether or not the certificate is obtained, the regret is bounded by that divergence. The margin depends on the family of admissible rules, and we show why this dependence is not a matter of notational convenience. The relevant quantity here is the divergence generated by the cost gaps between rules. This construction belongs to a lineage of task-oriented divergences; we specifically justify the choice of the class of loss differences by its invariance under the addition of a cost common to all rules and by its direct link with changes of trade-off. We then show that this class pulls back exactly along a pipeline: the divergence measured at the output is exactly the one generated, at the source, by the loss class pulled back along the transformations. Decision quality therefore does not propagate as an intrinsic score from upstream to downstream; the class that defines it is induced by the downstream use and pulled back upstream. The usual Lipschitz bounds then appear as computable relaxations of this exact identity — under regularity conditions that constrain the rules as well, and which we make explicit. We finally state five open problems.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Data Quality and Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.