From One Case to a Pattern: Convergence-Based Label Noise Detection Across Noise Types and Class Sizes

Reliable evaluation of machine learning systems depends on trustworthy labels, but existing label-noise detectors typically require either ground truth to validate against or expensive per-example computation to identify errors. We show that a simpler, aggregate signal is already available: when several class-balancing techniques are applied to the same dataset, their disagreement on target-class recall reliably tracks the amount of label noise present, offering evaluators a way to flag potentially unreliable training or evaluation labels without needing ground truth. We validate this signal rigorously, not just observationally: across random and systematic noise injection, three target-class sizes, and two model architectures, using the coefficient of variation (CV) as a metric corrected for two confounds identified during initial testing (a floor effect and outlier domination from one resampling method), and confirmed statistically through repeated-seed experiments (5 to 10 independent seeds per condition) with 95% confidence intervals throughout. The core trend holds robustly in five of six model-by-class-size combinations tested; the sixth, a neural network on a middle-sized class, reveals a genuine architecture-specific limit to the signal's reliability rather than a uniform guarantee. We further compare this aggregate signal against confident learning, an established per-example detector: both track injected noise reliably, but CV is substantially more computationally expensive, a structural consequence of requiring multiple model fits to measure disagreement rather than a single calibrated classifier. We identify a concrete technical setting where an aggregate signal is nonetheless preferable, classifiers without calibrated probability outputs, and report these findings, their architecture-dependent behavior, and open questions for developing this signal into a practical diagnostic tool as directions for future work.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23158886
Primary Topic
Machine Learning and Data Classification
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

From One Case to a Pattern: Convergence-Based Label Noise Detection Across Noise Types and Class Sizes

Mounisha Roy
Zenodo (CERN European Organization for Nuclear Research)
Machine Learning and Data Classification
preprint

From One Case to a Pattern: Convergence-Based Label Noise Detection Across Noise Types and Class Sizes

Mounisha Roy
preprint en

Abstract

Reliable evaluation of machine learning systems depends on trustworthy labels, but existing label-noise detectors typically require either ground truth to validate against or expensive per-example computation to identify errors. We show that a simpler, aggregate signal is already available: when several class-balancing techniques are applied to the same dataset, their disagreement on target-class recall reliably tracks the amount of label noise present, offering evaluators a way to flag potentially unreliable training or evaluation labels without needing ground truth. We validate this signal rigorously, not just observationally: across random and systematic noise injection, three target-class sizes, and two model architectures, using the coefficient of variation (CV) as a metric corrected for two confounds identified during initial testing (a floor effect and outlier domination from one resampling method), and confirmed statistically through repeated-seed experiments (5 to 10 independent seeds per condition) with 95% confidence intervals throughout. The core trend holds robustly in five of six model-by-class-size combinations tested; the sixth, a neural network on a middle-sized class, reveals a genuine architecture-specific limit to the signal's reliability rather than a uniform guarantee. We further compare this aggregate signal against confident learning, an established per-example detector: both track injected noise reliably, but CV is substantially more computationally expensive, a structural consequence of requiring multiple model fits to measure disagreement rather than a single calibrated classifier. We identify a concrete technical setting where an aggregate signal is nonetheless preferable, classifiers without calibrated probability outputs, and report these findings, their architecture-dependent behavior, and open questions for developing this signal into a practical diagnostic tool as directions for future work.

Zenodo (CERN European Organization for Nuclear Research)
Independent Research Association (RO)
Machine Learning and Data Classification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

From One Case to a Pattern: Convergence-Based Label Noise Detection Across Noise Types and Class Sizes — Mounisha Roy · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS