Comparative Evaluation of Machine Learning Models for Breast Cancer Diagnosis: A Reproducible Benchmark Study on the Breast Cancer Wisconsin (Diagnostic) Dataset
This record contains the revised manuscript and reproducibility materials for a comparative benchmark study of machine learning models for breast cancer diagnosis using the Breast Cancer Wisconsin (Diagnostic) dataset. The study evaluates Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, XGBoost, and a neural-network classifier using a stratified train-test split and stratified five-fold cross-validation. The package includes the revised manuscript, source code, computational environment specification, analysis results, figures, citation metadata, and reproducibility documentation. The dataset used in the study is the Breast Cancer Wisconsin (Diagnostic) dataset available from the UCI Machine Learning Repository (Dataset ID 17; DOI: 10.24432/C5DW2B). The analysis is intended as a reproducible machine-learning benchmark and does not constitute clinical validation or a clinical deployment recommendation.
Authors
- ANUJ KUMAR SAXENA
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-29
- DOI
- https://doi.org/10.5281/zenodo.23040018
- Primary Topic
- AI in cancer detection
- Type
- preprint