Diagnosing a Lung Nodule Classification Pipeline: Near-Duplicate Leakage and Representation Mismatch on IQ-OTH/NCCD
Reported accuracies on the IQ-OTH/NCCD lung cancer dataset commonly exceed 98%. We show that a substantial part of that performance is consistent with near-duplicate images spanning the train/test boundary, and we measure the effect directly rather than by inference from a change of protocol. Under the image-level partition conventional on this benchmark, 41% of the training set is a near-duplicate, at cosine similarity >= 0.98, of some test image. Deleting those images from training alone, holding the test folds byte-identical, costs 0.076 balanced accuracy for ResNet50 (95% CI 0.050 to 0.103) beyond an equally sized random deletion; for EfficientNet-B0 the same intervention at a looser 0.94 threshold costs 0.199 (0.147 to 0.251). Both survive correction for fold dependence and multiple testing. Partitioning instead by recovered similarity groups, and training a model under each protocol across five architectures and five folds, shows that the image-level convention inflates balanced accuracy by 24 to 37 points for every architecture that learns the task, in both convolutional and transformer families. A second construction of the grouped folds reproduces those magnitudes to within 4.1 points while changing which individual test survives correction for fold dependence and multiple testing, so we rest the claim on agreement in magnitude rather than on any single p-value. We also show that this style of grouping should not be called patient-level recovery. Scored against known identity on MSD Task06_Lung, our own threshold-selection rule returns 65 groups against 63 true cases, a near-exact count match, at a pairwise precision of 0.022. Agreement between a recovered group count and a documented case count is therefore close to no evidence that patients were recovered, which is the evidence such procedures typically rest on, ours included. Of eight studies on this benchmark obtained in full, seven do not report partitioning at the patient level, and the one that does cannot say how, since the public distribution carries no identifiers: the partition here is assertable but not verifiable. We separately show that the near-chance behaviour of a deployed segmentation-then-classification pipeline is not a deficiency of model capacity: the same weights score 0.902 on the distribution they were trained for. We make no claim about the correctness of any individual prior result.
Authors
- Haseeb (ORCID: https://orcid.org/0009-0002-9733-1633)
Institutions
- National University of Computer and Emerging Sciences (PK)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-24
- DOI
- https://doi.org/10.5281/zenodo.22932025
- Primary Topic
- Lung Cancer Diagnosis and Treatment
- Type
- preprint