Rows Are Not Samples: Duplicate Records, Evaluation Bias, and Ranking Distortion in Standard Tabular Benchmarks
Standard tabular benchmarks are routinely evaluated with random stratified K-fold cross-validation, a protocol that treats every row as an independent sample. When a dataset contains duplicate records, that assumption fails: the same record can appear in both the training and the test folds, and the test prediction can be obtained by lookup rather than by generalisation. We audit duplicate records across the union of two widely used OpenML benchmark suites and the classic UCI collection (161 datasets), characterise the resulting fold leakage, and quantify how much of the reported accuracy is an artefact of it. Of the 156 datasets we could audit, 75 (48.1%) contain at least one duplicated row, 52 contain at least 1% duplicated rows and 20 at least 30%; 35 datasets contain duplicate groups whose members carry different labels — an irreducible ambiguity that caps the accuracy attainable by any model that sees only the features, in the worst case at 0.499. Exposure is bounded by the duplication rate and fixed by the group sizes (a pair is exposed with probability 0.9 at K=10), so the leak of a split can be computed from a cheap audit rather than measured. In a protocol comparison over the same 82 datasets — the 68 whose duplication rate exceeds 0.5% plus 14 that contain no duplicated row at all as the null control — the published protocol is optimistic by +1.5 pp on average over 340 dataset–model pairs and by up to +40 pp on a single pair, most for the instance-based and tree families and least for the regularised one; it changes the leading model on a quarter of the duplicated datasets and reverses the significance of the top-1/top-2 comparison on 10 of the 68, whereas on the clean control arm the same contrast is -0.1 pp and changes no leader. We argue that the unit of reporting in tabular benchmarking should be the distinct record, not the row, and give a short reporting checklist plus the open audit script used here.
Authors
- Heng Li (ORCID: https://orcid.org/0000-0003-4874-2874)
Institutions
- Central South University (CN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23119064
- Primary Topic
- Software Engineering Research
- Type
- preprint