Reproducible and Explainable Machine Learning for Breast Cancer Classification: Sensitivity-Oriented Thresholding and Independent Methodological Replication

Background: WDBC is a small historical benchmark, and near-ceiling performance alone provides limited evidence of transportability. Methods: We evaluated five model families for discrimination, calibration, and paired statistical testing. We utilized sensitivity-oriented out-of-fold (OOF) thresholds and decision curve analysis (DCA) and conducted an independent TOMPEI-CMMD methodological replication (larger two-center mammography cohort with biopsy-confirmed diagnoses; final cohort: 1358 patients, 1380 breasts; held-out: 272 patients, 279 breasts). Results: Across 20 additional stratified WDBC partitions, ROC-AUC remained consistently high (model means 0.988–0.995), but the criterion-specific nominal leader changed across partitions. Held-out TOMPEI ROC-AUC ranged from 0.775 to 0.804. No model demonstrated statistically significant superiority in ROC-AUC or frozen-threshold balanced accuracy after patient-cluster paired bootstrap comparison and Holm correction. The independent methodological replication yielded more moderate discrimination than the WDBC benchmark. Frozen sensitivity-oriented OOF thresholds transported reasonably but not uniformly. Conclusions: Multi-dimensional evaluation is more defensible than nominal AUC ranking. Neither the WDBC benchmark results nor the TOMPEI methodological replication establish population screening performance or clinical deployment readiness; further prospective clinically representative validation is required.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-09-15
DOI
https://doi.org/10.3390/app16189164
Primary Topic
AI in cancer detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Reproducible and Explainable Machine Learning for Breast Cancer Classification: Sensitivity-Oriented Thresholding and Independent Methodological Replication

Younes Nadir, Abdellah Bakhouyi, Lahcen Amhaimar, Abderrahim Khalidi et al.
Applied Sciences
AI in cancer detection
article

Reproducible and Explainable Machine Learning for Breast Cancer Classification: Sensitivity-Oriented Thresholding and Independent Methodological Replication

Younes Nadir, Abdellah Bakhouyi, Lahcen Amhaimar, Abderrahim Khalidi, Mohamed Azzouazi, Mohamed Rachdi
article en

Abstract

Background: WDBC is a small historical benchmark, and near-ceiling performance alone provides limited evidence of transportability. Methods: We evaluated five model families for discrimination, calibration, and paired statistical testing. We utilized sensitivity-oriented out-of-fold (OOF) thresholds and decision curve analysis (DCA) and conducted an independent TOMPEI-CMMD methodological replication (larger two-center mammography cohort with biopsy-confirmed diagnoses; final cohort: 1358 patients, 1380 breasts; held-out: 272 patients, 279 breasts). Results: Across 20 additional stratified WDBC partitions, ROC-AUC remained consistently high (model means 0.988–0.995), but the criterion-specific nominal leader changed across partitions. Held-out TOMPEI ROC-AUC ranged from 0.775 to 0.804. No model demonstrated statistically significant superiority in ROC-AUC or frozen-threshold balanced accuracy after patient-cluster paired bootstrap comparison and Holm correction. The independent methodological replication yielded more moderate discrimination than the WDBC benchmark. Frozen sensitivity-oriented OOF thresholds transported reasonably but not uniformly. Conclusions: Multi-dimensional evaluation is more defensible than nominal AUC ranking. Neither the WDBC benchmark results nor the TOMPEI methodological replication establish population screening performance or clinical deployment readiness; further prospective clinically representative validation is required.

Applied SciencesVol. 16(18)
École Normale Supérieure de l'Enseignement Technique de Mohammedia (MA), Université Hassan II Mohammedia (MA), University of Hassan II Casablanca (MA)
Gender equality
Openalex Percentile: Top 8%
AI in cancer detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.