Rows Are Not Samples: Duplicate Records, Evaluation Bias, and Ranking Distortion in Standard Tabular Benchmarks

Standard tabular benchmarks are routinely evaluated with random stratified K-fold cross-validation, a protocol that treats every row as an independent sample. When a dataset contains duplicate records, that assumption fails: the same record can appear in both the training and the test folds, and the test prediction can be obtained by lookup rather than by generalisation. We audit duplicate records across the union of two widely used OpenML benchmark suites and the classic UCI collection (161 datasets), characterise the resulting fold leakage, and quantify how much of the reported accuracy is an artefact of it. Of the 156 datasets we could audit, 75 (48.1%) contain at least one duplicated row, 52 contain at least 1% duplicated rows and 20 at least 30%; 35 datasets contain duplicate groups whose members carry different labels — an irreducible ambiguity that caps the accuracy attainable by any model that sees only the features, in the worst case at 0.499. Exposure is bounded by the duplication rate and fixed by the group sizes (a pair is exposed with probability 0.9 at K=10), so the leak of a split can be computed from a cheap audit rather than measured. In a protocol comparison over the same 82 datasets — the 68 whose duplication rate exceeds 0.5% plus 14 that contain no duplicated row at all as the null control — the published protocol is optimistic by +1.5 pp on average over 340 dataset–model pairs and by up to +40 pp on a single pair, most for the instance-based and tree families and least for the regularised one; it changes the leading model on a quarter of the duplicated datasets and reverses the significance of the top-1/top-2 comparison on 10 of the 68, whereas on the clean control arm the same contrast is -0.1 pp and changes no leader. We argue that the unit of reporting in tabular benchmarking should be the distinct record, not the row, and give a short reporting checklist plus the open audit script used here.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23119064
Primary Topic
Software Engineering Research
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Rows Are Not Samples: Duplicate Records, Evaluation Bias, and Ranking Distortion in Standard Tabular Benchmarks

Heng Li
Zenodo (CERN European Organization for Nuclear Research)
Software Engineering Research
preprint

Rows Are Not Samples: Duplicate Records, Evaluation Bias, and Ranking Distortion in Standard Tabular Benchmarks

Heng Li
preprint en

Abstract

Standard tabular benchmarks are routinely evaluated with random stratified K-fold cross-validation, a protocol that treats every row as an independent sample. When a dataset contains duplicate records, that assumption fails: the same record can appear in both the training and the test folds, and the test prediction can be obtained by lookup rather than by generalisation. We audit duplicate records across the union of two widely used OpenML benchmark suites and the classic UCI collection (161 datasets), characterise the resulting fold leakage, and quantify how much of the reported accuracy is an artefact of it. Of the 156 datasets we could audit, 75 (48.1%) contain at least one duplicated row, 52 contain at least 1% duplicated rows and 20 at least 30%; 35 datasets contain duplicate groups whose members carry different labels — an irreducible ambiguity that caps the accuracy attainable by any model that sees only the features, in the worst case at 0.499. Exposure is bounded by the duplication rate and fixed by the group sizes (a pair is exposed with probability 0.9 at K=10), so the leak of a split can be computed from a cheap audit rather than measured. In a protocol comparison over the same 82 datasets — the 68 whose duplication rate exceeds 0.5% plus 14 that contain no duplicated row at all as the null control — the published protocol is optimistic by +1.5 pp on average over 340 dataset–model pairs and by up to +40 pp on a single pair, most for the instance-based and tree families and least for the regularised one; it changes the leading model on a quarter of the duplicated datasets and reverses the significance of the top-1/top-2 comparison on 10 of the 68, whereas on the clean control arm the same contrast is -0.1 pp and changes no leader. We argue that the unit of reporting in tabular benchmarking should be the distinct record, not the row, and give a short reporting checklist plus the open audit script used here.

Zenodo (CERN European Organization for Nuclear Research)
Central South University (CN)
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Rows Are Not Samples: Duplicate Records, Evaluation Bias, and Ranking Distortion in Standard Tabular Benchmarks — Heng Li · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS