Machine learning for predicting mortality in women diagnosed with breast cancer in the State of Mato Grosso, Brazil using linked population-based cancer registry and mortality data
BACKGROUND: Breast cancer is the most frequently diagnosed malignancy and a leading cause of cancer-related mortality among women worldwide. In Brazil, pronounced regional inequalities persist in access to timely diagnosis and treatment, particularly in large and socioeconomically heterogeneous states such as Mato Grosso. This study aimed to develop and validate machine learning (ML) models to predict 1-, 3-, 5-, and 10-year all-cause and breast cancer-specific mortality among women diagnosed with breast cancer in Mato Grosso, Brazil. METHODS: A retrospective, population-based cohort study was conducted, including 7,815 women diagnosed with breast cancer between 2001 and 2018. The dataset was randomly split into training (75%) and testing (25%) sets. Logistic regression, Random Forest, XGBoost, CatBoost, and LightGBM models were developed to predict all-cause and breast cancer-specific mortality at 1, 3, 5, and 10 years. Model performance was primarily evaluated using the area under the receiver operating characteristic curve (AUROC). Calibration was assessed using calibration plots and the Brier score (BS), with 95% confidence intervals estimated via bootstrap resampling. Model interpretability was evaluated using Shapley Additive Explanations (SHAP). RESULTS: Gradient boosting models consistently demonstrated superior performance across prediction horizons. For all-cause mortality, CatBoost achieved the highest discrimination at 1 year (AUROC 83.18%, 95% CI 80.01-86.33), while LightGBM showed higher discrimination at 3 and 5 years (AUROC 81.01%, 95% CI 78.63-83.42; and AUROC 78.39%, 95% CI 75.99-80.64, respectively). For breast cancer-specific mortality, logistic regression achieved the highest AUROC at 1 year (84.44%, 95% CI 81.04-87.35), whereas LightGBM achieved the highest at 3 years (AUROC 79.82%, 95% CI 77.15-82.34), XGBoost at 5 years (AUROC 80.79%, 95% CI 78.54-83.23). SHAP analyses consistently identified age at diagnosis, metastatic stage, histological diagnosis, invasive ductal carcinoma (IDC), and marital status as the most influential predictors. CONCLUSIONS: ML models demonstrated good and stable discriminative performance in predicting breast cancer mortality using routine population-based data. These models may support the early identification of high-risk patients and help guide risk stratification and follow-up planning in resource-constrained health systems. However, the models have not yet been externally validated, and their generalizability to populations beyond Mato Grosso remains uncertain. External validation in independent populations is therefore warranted before broader implementation in other regions of Brazil.
Authors
- Audêncio Victor (ORCID: https://orcid.org/0000-0002-8161-3639)
- Sancho Pedro Xavier (ORCID: https://orcid.org/0000-0001-9493-4098)
- Manuel Mahoche (ORCID: https://orcid.org/0000-0002-9784-6402)
- Marco Aurélio Bertúlio das Neves (ORCID: https://orcid.org/0000-0002-0685-9233)
- Ana Raquel Manuel Gotine (ORCID: https://orcid.org/0000-0002-3539-4236)
- Noemi Dreyer Galvão (ORCID: https://orcid.org/0000-0002-8337-0669)
- Alexandre Dias Porto Chiavegatto Filho (ORCID: https://orcid.org/0000-0003-3251-9600)
- Fernando Henrique de Albuquerque Maia (ORCID: https://orcid.org/0000-0001-7227-9774)
- Ageo Mario Cândido da Silva
Institutions
- Universidade de São Paulo (BR)
- Universidade do Estado de Mato Grosso (BR)
- Universidade Federal de Mato Grosso (BR)
- London School of Hygiene & Tropical Medicine (GB)
Publication Details
- Journal
- PLoS ONE
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1371/journal.pone.0356621
- Primary Topic
- Global Cancer Incidence and Screening
- Type
- article
- Field-Weighted Citation Impact
- 0.00