Explainable Machine Learning for Oncological Mass Classification: Benchmarking, Threshold Optimization and Robustness Analysis
Breast cancer remains one of the most commonly diagnosed malignancies worldwide, and timely, accurate discrimination between benign and malignant breast masses is critical to reducing unnecessary biopsies while ensuring early treatment of true malignancies. This study presents a comprehensive, methodologically rigorous machine learning (ML) analysis of six supervised classifiers — Logistic Regression, Support Vector Machine (SVM), Random Forest, Gradient Boosting, k-Nearest Neighbors (k-NN), and a shallow Artificial Neural Network (ANN) — for binary classification of breast masses using the Wisconsin Diagnostic Breast Cancer (WDBC) dataset, a well-established, publicly available benchmark comprising 569 cases and 30 real-valued nuclear morphometric features computed from digitized fine-needle aspirate (FNA) images. Beyond standard accuracy benchmarking, this study contributes four analyses that are comparatively underreported in the WDBC literature: a feature-domain ablation quantifying the diagnostic contribution of mean-value, standard-error, and worst-value feature subsets; a class-imbalance handling comparison; a clinically motivated decision-threshold optimization using the Youden index; and a dual global-interpretability analysis combining Random Forest Gini importance with SHapley Additive exPlanations (SHAP). Each classifier was tuned via 5-fold cross-validated grid search and evaluated on a stratified 75/25 held-out split. Logistic Regression and k-Nearest Neighbors jointly achieved the highest held-out test accuracy (97.90%), with all six classifiers exceeding 95.8% accuracy; pairwise paired t-tests over cross-validation folds found only one statistically significant difference among fifteen pairwise comparisons, indicating that classifier choice for this task is largely accuracy-neutral. The best-performing model achieved a receiver operating characteristic area under the curve (ROC-AUC) of 0.997, an average precision of 0.998, and a Matthews correlation coefficient of 0.955. Feature-domain ablation showed that worst-value features alone recovered 96.50% accuracy versus 97.90% for the full feature set, while standard-error features alone achieved only 86.71%, and SHAP analysis corroborated Random Forest's Gini-based ranking, jointly identifying worst area, worst perimeter, and worst/mean concave points as the dominant diagnostic drivers. Youden-index threshold optimization recovered two additional correctly identified malignant cases relative to the default 0.5 probability threshold, at a small, quantified cost to specificity. These results, obtained on real diagnostic data rather than simulated signals, corroborate and extend a substantial body of prior work, and reinforce the broader case for feature-based, interpretable, computationally lightweight ML as a viable decision-support tool for cytological breast mass classification, while underscoring that any such tool requires prospective, multi-institutional clinical validation before informing real diagnostic decisions. Keywords: machine learning; explainable AI; SHAP; breast cancer; Computer-Aided Diagnosis.; diagnostic classification; Wisconsin Diagnostic Breast Cancer dataset; fine-needle aspiration; decision threshold optimization; feature ablation
Authors
- Rajesh K. Mishra (ORCID: https://orcid.org/0000-0002-6504-5766)
- Divyansh Mishra
- Rekha Agarwal (ORCID: https://orcid.org/0009-0008-9451-3136)
Institutions
- Tropical Forest Research Institute (IN)
- Rani Durgavati University (IN)
- Xavier School of Management (IN)
Publication Details
- Journal
- International Journal of Oncology Research
- Published
- 2026-09-14
- DOI
- https://doi.org/10.64823/ijor.2601002
- Primary Topic
- AI in cancer detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00