Development of a machine learning model for risk stratification of cervical cancer based on cross-sectional data

Abstract Background Cervical cancer ranks as the fourth most common malignancy among women worldwide, with particularly poor prognosis and high rates of late-stage diagnosis in low- and middle-income countries. The current FIGO staging system has limited prognostic precision at the individual-level risk assessment, underscoring the need for diagnostic tools to support timely clinical decision-making. Methods This study aimed to develop and validate a machine learning model that combines clinical, molecular, and cytological features for individualized risk stratification of cervical cancer at a single point in time. A total of 207 participants from the Fifth Affiliated Hospital of Sun Yat-sen University were enrolled, including 151 non-cancer and 56 cancer cases. Over 20 variables were collected, including age, blood pressure, HPV genotypes, P16 expression, serum tumor markers (e.g., SCC, CEA, HE4), and cytological grading. Feature selection was performed using LASSO regression-based feature ranking, and the optimal number of predictors was determined by evaluating Random Forest (RF) performance across different feature subsets, resulting in 11 core predictors (e.g., SCC, HPV16/18, HSIL). Results Eleven machine learning models were evaluated. RF and LR demonstrated the strongest overall performance, with RF achieving the highest AUROC (0.966) and sensitivity (100%) among the evaluated models. RF was subsequently selected as the final model based on its balanced discrimination and calibration performance. SHAP analysis identified SCC, HPV16/18, and documented P16 status as major contributors. A nomogram was constructed for clinical visualization and interpretation. Conclusions This study presents an interpretable machine learning model for cross-sectional risk stratification of cervical cancer, providing a potential approach to support individualized screening and diagnostic decision-making. Further multicenter validation and integration of longitudinal data are warranted to enhance generalizability and future prognostic use. We acknowledge that the current findings are based on single-center data, and external validation is a key next step.

Authors

Institutions

Publication Details

Journal
BMC Medical Genomics
Published
2026-09-24
DOI
https://doi.org/10.1186/s12920-026-02483-7
Primary Topic
Cervical Cancer and HPV Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Development of a machine learning model for risk stratification of cervical cancer based on cross-sectional data

Bingfan Xie, Shaoxia Zhang, Feng Wang, Dandan Huang et al.
BMC Medical Genomics
Cervical Cancer and HPV Research
article

Development of a machine learning model for risk stratification of cervical cancer based on cross-sectional data

Bingfan Xie, Shaoxia Zhang, Feng Wang, Dandan Huang, Lili Gu, Honghui Ou, Cuiting Hu, Chunrong Qin, Yong Chen, Qizhao Nie
article en

Abstract

Abstract Background Cervical cancer ranks as the fourth most common malignancy among women worldwide, with particularly poor prognosis and high rates of late-stage diagnosis in low- and middle-income countries. The current FIGO staging system has limited prognostic precision at the individual-level risk assessment, underscoring the need for diagnostic tools to support timely clinical decision-making. Methods This study aimed to develop and validate a machine learning model that combines clinical, molecular, and cytological features for individualized risk stratification of cervical cancer at a single point in time. A total of 207 participants from the Fifth Affiliated Hospital of Sun Yat-sen University were enrolled, including 151 non-cancer and 56 cancer cases. Over 20 variables were collected, including age, blood pressure, HPV genotypes, P16 expression, serum tumor markers (e.g., SCC, CEA, HE4), and cytological grading. Feature selection was performed using LASSO regression-based feature ranking, and the optimal number of predictors was determined by evaluating Random Forest (RF) performance across different feature subsets, resulting in 11 core predictors (e.g., SCC, HPV16/18, HSIL). Results Eleven machine learning models were evaluated. RF and LR demonstrated the strongest overall performance, with RF achieving the highest AUROC (0.966) and sensitivity (100%) among the evaluated models. RF was subsequently selected as the final model based on its balanced discrimination and calibration performance. SHAP analysis identified SCC, HPV16/18, and documented P16 status as major contributors. A nomogram was constructed for clinical visualization and interpretation. Conclusions This study presents an interpretable machine learning model for cross-sectional risk stratification of cervical cancer, providing a potential approach to support individualized screening and diagnostic decision-making. Further multicenter validation and integration of longitudinal data are warranted to enhance generalizability and future prognostic use. We acknowledge that the current findings are based on single-center data, and external validation is a key next step.

BMC Medical Genomics
Sun Yat-sen University (CN), Shenzhen Maternity and Child Healthcare Hospital (CN), Fifth Affiliated Hospital of Sun Yat-sen University (CN), Anyang Institute of Technology (CN)
Openalex Percentile: Top 11%
Cervical Cancer and HPV Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.