Machine learning for predicting mortality in women diagnosed with breast cancer in the State of Mato Grosso, Brazil using linked population-based cancer registry and mortality data

BACKGROUND: Breast cancer is the most frequently diagnosed malignancy and a leading cause of cancer-related mortality among women worldwide. In Brazil, pronounced regional inequalities persist in access to timely diagnosis and treatment, particularly in large and socioeconomically heterogeneous states such as Mato Grosso. This study aimed to develop and validate machine learning (ML) models to predict 1-, 3-, 5-, and 10-year all-cause and breast cancer-specific mortality among women diagnosed with breast cancer in Mato Grosso, Brazil. METHODS: A retrospective, population-based cohort study was conducted, including 7,815 women diagnosed with breast cancer between 2001 and 2018. The dataset was randomly split into training (75%) and testing (25%) sets. Logistic regression, Random Forest, XGBoost, CatBoost, and LightGBM models were developed to predict all-cause and breast cancer-specific mortality at 1, 3, 5, and 10 years. Model performance was primarily evaluated using the area under the receiver operating characteristic curve (AUROC). Calibration was assessed using calibration plots and the Brier score (BS), with 95% confidence intervals estimated via bootstrap resampling. Model interpretability was evaluated using Shapley Additive Explanations (SHAP). RESULTS: Gradient boosting models consistently demonstrated superior performance across prediction horizons. For all-cause mortality, CatBoost achieved the highest discrimination at 1 year (AUROC 83.18%, 95% CI 80.01-86.33), while LightGBM showed higher discrimination at 3 and 5 years (AUROC 81.01%, 95% CI 78.63-83.42; and AUROC 78.39%, 95% CI 75.99-80.64, respectively). For breast cancer-specific mortality, logistic regression achieved the highest AUROC at 1 year (84.44%, 95% CI 81.04-87.35), whereas LightGBM achieved the highest at 3 years (AUROC 79.82%, 95% CI 77.15-82.34), XGBoost at 5 years (AUROC 80.79%, 95% CI 78.54-83.23). SHAP analyses consistently identified age at diagnosis, metastatic stage, histological diagnosis, invasive ductal carcinoma (IDC), and marital status as the most influential predictors. CONCLUSIONS: ML models demonstrated good and stable discriminative performance in predicting breast cancer mortality using routine population-based data. These models may support the early identification of high-risk patients and help guide risk stratification and follow-up planning in resource-constrained health systems. However, the models have not yet been externally validated, and their generalizability to populations beyond Mato Grosso remains uncertain. External validation in independent populations is therefore warranted before broader implementation in other regions of Brazil.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-18
DOI
https://doi.org/10.1371/journal.pone.0356621
Primary Topic
Global Cancer Incidence and Screening
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Machine learning for predicting mortality in women diagnosed with breast cancer in the State of Mato Grosso, Brazil using linked population-based cancer registry and mortality data

Audêncio Victor, Sancho Pedro Xavier, Manuel Mahoche, Marco Aurélio Bertúlio das Neves et al.
PLoS ONE
Global Cancer Incidence and Screening
article

Machine learning for predicting mortality in women diagnosed with breast cancer in the State of Mato Grosso, Brazil using linked population-based cancer registry and mortality data

Audêncio Victor, Sancho Pedro Xavier, Manuel Mahoche, Marco Aurélio Bertúlio das Neves, Ana Raquel Manuel Gotine, Noemi Dreyer Galvão, Alexandre Dias Porto Chiavegatto Filho, Fernando Henrique de Albuquerque Maia, Ageo Mario Cândido da Silva
article en

Abstract

BACKGROUND: Breast cancer is the most frequently diagnosed malignancy and a leading cause of cancer-related mortality among women worldwide. In Brazil, pronounced regional inequalities persist in access to timely diagnosis and treatment, particularly in large and socioeconomically heterogeneous states such as Mato Grosso. This study aimed to develop and validate machine learning (ML) models to predict 1-, 3-, 5-, and 10-year all-cause and breast cancer-specific mortality among women diagnosed with breast cancer in Mato Grosso, Brazil. METHODS: A retrospective, population-based cohort study was conducted, including 7,815 women diagnosed with breast cancer between 2001 and 2018. The dataset was randomly split into training (75%) and testing (25%) sets. Logistic regression, Random Forest, XGBoost, CatBoost, and LightGBM models were developed to predict all-cause and breast cancer-specific mortality at 1, 3, 5, and 10 years. Model performance was primarily evaluated using the area under the receiver operating characteristic curve (AUROC). Calibration was assessed using calibration plots and the Brier score (BS), with 95% confidence intervals estimated via bootstrap resampling. Model interpretability was evaluated using Shapley Additive Explanations (SHAP). RESULTS: Gradient boosting models consistently demonstrated superior performance across prediction horizons. For all-cause mortality, CatBoost achieved the highest discrimination at 1 year (AUROC 83.18%, 95% CI 80.01-86.33), while LightGBM showed higher discrimination at 3 and 5 years (AUROC 81.01%, 95% CI 78.63-83.42; and AUROC 78.39%, 95% CI 75.99-80.64, respectively). For breast cancer-specific mortality, logistic regression achieved the highest AUROC at 1 year (84.44%, 95% CI 81.04-87.35), whereas LightGBM achieved the highest at 3 years (AUROC 79.82%, 95% CI 77.15-82.34), XGBoost at 5 years (AUROC 80.79%, 95% CI 78.54-83.23). SHAP analyses consistently identified age at diagnosis, metastatic stage, histological diagnosis, invasive ductal carcinoma (IDC), and marital status as the most influential predictors. CONCLUSIONS: ML models demonstrated good and stable discriminative performance in predicting breast cancer mortality using routine population-based data. These models may support the early identification of high-risk patients and help guide risk stratification and follow-up planning in resource-constrained health systems. However, the models have not yet been externally validated, and their generalizability to populations beyond Mato Grosso remains uncertain. External validation in independent populations is therefore warranted before broader implementation in other regions of Brazil.

PLoS ONEVol. 21(9)
Universidade de São Paulo (BR), Universidade do Estado de Mato Grosso (BR), Universidade Federal de Mato Grosso (BR), London School of Hygiene & Tropical Medicine (GB)
Gender equality
Openalex Percentile: Top 14%
Global Cancer Incidence and Screening
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.