Diagnosing a Lung Nodule Classification Pipeline: Near-Duplicate Leakage and Representation Mismatch on IQ-OTH/NCCD

Reported accuracies on the IQ-OTH/NCCD lung cancer dataset commonly exceed 98%. We show that a substantial part of that performance is consistent with near-duplicate images spanning the train/test boundary, and we measure the effect directly rather than by inference from a change of protocol. Under the image-level partition conventional on this benchmark, 41% of the training set is a near-duplicate, at cosine similarity >= 0.98, of some test image. Deleting those images from training alone, holding the test folds byte-identical, costs 0.076 balanced accuracy for ResNet50 (95% CI 0.050 to 0.103) beyond an equally sized random deletion; for EfficientNet-B0 the same intervention at a looser 0.94 threshold costs 0.199 (0.147 to 0.251). Both survive correction for fold dependence and multiple testing. Partitioning instead by recovered similarity groups, and training a model under each protocol across five architectures and five folds, shows that the image-level convention inflates balanced accuracy by 24 to 37 points for every architecture that learns the task, in both convolutional and transformer families. A second construction of the grouped folds reproduces those magnitudes to within 4.1 points while changing which individual test survives correction for fold dependence and multiple testing, so we rest the claim on agreement in magnitude rather than on any single p-value. We also show that this style of grouping should not be called patient-level recovery. Scored against known identity on MSD Task06_Lung, our own threshold-selection rule returns 65 groups against 63 true cases, a near-exact count match, at a pairwise precision of 0.022. Agreement between a recovered group count and a documented case count is therefore close to no evidence that patients were recovered, which is the evidence such procedures typically rest on, ours included. Of eight studies on this benchmark obtained in full, seven do not report partitioning at the patient level, and the one that does cannot say how, since the public distribution carries no identifiers: the partition here is assertable but not verifiable. We separately show that the near-chance behaviour of a deployed segmentation-then-classification pipeline is not a deficiency of model capacity: the same weights score 0.902 on the distribution they were trained for. We make no claim about the correctness of any individual prior result.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-24
DOI
https://doi.org/10.5281/zenodo.22932025
Primary Topic
Lung Cancer Diagnosis and Treatment
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Diagnosing a Lung Nodule Classification Pipeline: Near-Duplicate Leakage and Representation Mismatch on IQ-OTH/NCCD

Haseeb
Zenodo (CERN European Organization for Nuclear Research)
Lung Cancer Diagnosis and Treatment
preprint

Diagnosing a Lung Nodule Classification Pipeline: Near-Duplicate Leakage and Representation Mismatch on IQ-OTH/NCCD

Haseeb
preprint en

Abstract

Reported accuracies on the IQ-OTH/NCCD lung cancer dataset commonly exceed 98%. We show that a substantial part of that performance is consistent with near-duplicate images spanning the train/test boundary, and we measure the effect directly rather than by inference from a change of protocol. Under the image-level partition conventional on this benchmark, 41% of the training set is a near-duplicate, at cosine similarity >= 0.98, of some test image. Deleting those images from training alone, holding the test folds byte-identical, costs 0.076 balanced accuracy for ResNet50 (95% CI 0.050 to 0.103) beyond an equally sized random deletion; for EfficientNet-B0 the same intervention at a looser 0.94 threshold costs 0.199 (0.147 to 0.251). Both survive correction for fold dependence and multiple testing. Partitioning instead by recovered similarity groups, and training a model under each protocol across five architectures and five folds, shows that the image-level convention inflates balanced accuracy by 24 to 37 points for every architecture that learns the task, in both convolutional and transformer families. A second construction of the grouped folds reproduces those magnitudes to within 4.1 points while changing which individual test survives correction for fold dependence and multiple testing, so we rest the claim on agreement in magnitude rather than on any single p-value. We also show that this style of grouping should not be called patient-level recovery. Scored against known identity on MSD Task06_Lung, our own threshold-selection rule returns 65 groups against 63 true cases, a near-exact count match, at a pairwise precision of 0.022. Agreement between a recovered group count and a documented case count is therefore close to no evidence that patients were recovered, which is the evidence such procedures typically rest on, ours included. Of eight studies on this benchmark obtained in full, seven do not report partitioning at the patient level, and the one that does cannot say how, since the public distribution carries no identifiers: the partition here is assertable but not verifiable. We separately show that the near-chance behaviour of a deployed segmentation-then-classification pipeline is not a deficiency of model capacity: the same weights score 0.902 on the distribution they were trained for. We make no claim about the correctness of any individual prior result.

Zenodo (CERN European Organization for Nuclear Research)
National University of Computer and Emerging Sciences (PK)
Lung Cancer Diagnosis and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.