Explainable Machine Learning for Oncological Mass Classification: Benchmarking, Threshold Optimization and Robustness Analysis

Breast cancer remains one of the most commonly diagnosed malignancies worldwide, and timely, accurate discrimination between benign and malignant breast masses is critical to reducing unnecessary biopsies while ensuring early treatment of true malignancies. This study presents a comprehensive, methodologically rigorous machine learning (ML) analysis of six supervised classifiers — Logistic Regression, Support Vector Machine (SVM), Random Forest, Gradient Boosting, k-Nearest Neighbors (k-NN), and a shallow Artificial Neural Network (ANN) — for binary classification of breast masses using the Wisconsin Diagnostic Breast Cancer (WDBC) dataset, a well-established, publicly available benchmark comprising 569 cases and 30 real-valued nuclear morphometric features computed from digitized fine-needle aspirate (FNA) images. Beyond standard accuracy benchmarking, this study contributes four analyses that are comparatively underreported in the WDBC literature: a feature-domain ablation quantifying the diagnostic contribution of mean-value, standard-error, and worst-value feature subsets; a class-imbalance handling comparison; a clinically motivated decision-threshold optimization using the Youden index; and a dual global-interpretability analysis combining Random Forest Gini importance with SHapley Additive exPlanations (SHAP). Each classifier was tuned via 5-fold cross-validated grid search and evaluated on a stratified 75/25 held-out split. Logistic Regression and k-Nearest Neighbors jointly achieved the highest held-out test accuracy (97.90%), with all six classifiers exceeding 95.8% accuracy; pairwise paired t-tests over cross-validation folds found only one statistically significant difference among fifteen pairwise comparisons, indicating that classifier choice for this task is largely accuracy-neutral. The best-performing model achieved a receiver operating characteristic area under the curve (ROC-AUC) of 0.997, an average precision of 0.998, and a Matthews correlation coefficient of 0.955. Feature-domain ablation showed that worst-value features alone recovered 96.50% accuracy versus 97.90% for the full feature set, while standard-error features alone achieved only 86.71%, and SHAP analysis corroborated Random Forest's Gini-based ranking, jointly identifying worst area, worst perimeter, and worst/mean concave points as the dominant diagnostic drivers. Youden-index threshold optimization recovered two additional correctly identified malignant cases relative to the default 0.5 probability threshold, at a small, quantified cost to specificity. These results, obtained on real diagnostic data rather than simulated signals, corroborate and extend a substantial body of prior work, and reinforce the broader case for feature-based, interpretable, computationally lightweight ML as a viable decision-support tool for cytological breast mass classification, while underscoring that any such tool requires prospective, multi-institutional clinical validation before informing real diagnostic decisions. Keywords: machine learning; explainable AI; SHAP; breast cancer; Computer-Aided Diagnosis.; diagnostic classification; Wisconsin Diagnostic Breast Cancer dataset; fine-needle aspiration; decision threshold optimization; feature ablation

Authors

Institutions

Publication Details

Journal
International Journal of Oncology Research
Published
2026-09-14
DOI
https://doi.org/10.64823/ijor.2601002
Primary Topic
AI in cancer detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Explainable Machine Learning for Oncological Mass Classification: Benchmarking, Threshold Optimization and Robustness Analysis

Rajesh K. Mishra, Divyansh Mishra, Rekha Agarwal
International Journal of Oncology Research
AI in cancer detection
article

Explainable Machine Learning for Oncological Mass Classification: Benchmarking, Threshold Optimization and Robustness Analysis

Rajesh K. Mishra, Divyansh Mishra, Rekha Agarwal
article en

Abstract

Breast cancer remains one of the most commonly diagnosed malignancies worldwide, and timely, accurate discrimination between benign and malignant breast masses is critical to reducing unnecessary biopsies while ensuring early treatment of true malignancies. This study presents a comprehensive, methodologically rigorous machine learning (ML) analysis of six supervised classifiers — Logistic Regression, Support Vector Machine (SVM), Random Forest, Gradient Boosting, k-Nearest Neighbors (k-NN), and a shallow Artificial Neural Network (ANN) — for binary classification of breast masses using the Wisconsin Diagnostic Breast Cancer (WDBC) dataset, a well-established, publicly available benchmark comprising 569 cases and 30 real-valued nuclear morphometric features computed from digitized fine-needle aspirate (FNA) images. Beyond standard accuracy benchmarking, this study contributes four analyses that are comparatively underreported in the WDBC literature: a feature-domain ablation quantifying the diagnostic contribution of mean-value, standard-error, and worst-value feature subsets; a class-imbalance handling comparison; a clinically motivated decision-threshold optimization using the Youden index; and a dual global-interpretability analysis combining Random Forest Gini importance with SHapley Additive exPlanations (SHAP). Each classifier was tuned via 5-fold cross-validated grid search and evaluated on a stratified 75/25 held-out split. Logistic Regression and k-Nearest Neighbors jointly achieved the highest held-out test accuracy (97.90%), with all six classifiers exceeding 95.8% accuracy; pairwise paired t-tests over cross-validation folds found only one statistically significant difference among fifteen pairwise comparisons, indicating that classifier choice for this task is largely accuracy-neutral. The best-performing model achieved a receiver operating characteristic area under the curve (ROC-AUC) of 0.997, an average precision of 0.998, and a Matthews correlation coefficient of 0.955. Feature-domain ablation showed that worst-value features alone recovered 96.50% accuracy versus 97.90% for the full feature set, while standard-error features alone achieved only 86.71%, and SHAP analysis corroborated Random Forest's Gini-based ranking, jointly identifying worst area, worst perimeter, and worst/mean concave points as the dominant diagnostic drivers. Youden-index threshold optimization recovered two additional correctly identified malignant cases relative to the default 0.5 probability threshold, at a small, quantified cost to specificity. These results, obtained on real diagnostic data rather than simulated signals, corroborate and extend a substantial body of prior work, and reinforce the broader case for feature-based, interpretable, computationally lightweight ML as a viable decision-support tool for cytological breast mass classification, while underscoring that any such tool requires prospective, multi-institutional clinical validation before informing real diagnostic decisions. Keywords: machine learning; explainable AI; SHAP; breast cancer; Computer-Aided Diagnosis.; diagnostic classification; Wisconsin Diagnostic Breast Cancer dataset; fine-needle aspiration; decision threshold optimization; feature ablation

International Journal of Oncology ResearchVol. 1(1)
Tropical Forest Research Institute (IN), Rani Durgavati University (IN), Xavier School of Management (IN)
Openalex Percentile: Top 8%
AI in cancer detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.