CT venous-phase radiomics for preoperative differentiation of early and advanced gastric cancer: construction, comparison, and validation of machine-learning models

Abstract Objective To construct various radiomics models based on CT venous-phase images for the quantitative assessment of gastric cancer (GC) invasion depth into the gastric wall, and to compare the diagnostic performance of different machine-learning models in distinguishing early (T1–T2 stage) from advanced (T3–T4 stage) GC, aiming to identify a robust and interpretable model and evaluate its potential for clinical application. Methods In this retrospective study, 223 pathologically confirmed GC patients (66 early-stage, 157 advanced-stage) treated between January 2022 and May 2025 were enrolled. Patients were allocated into a training set ( n = 156) and an internal hold-out test set ( n = 67) through stratified random sampling at a 7:3 ratio. All patients underwent enhanced CT within one week prior to surgery. Venous-phase images were selected, and three-dimensional regions of interest (ROIs) encompassing the tumor were manually delineated using 3D Slicer for radiomics feature extraction. To rigorously exclude data leakage, all preprocessing steps—including feature standardization, LASSO feature selection, hyperparameter tuning, and threshold determination—were performed exclusively within the training set using a pipeline-based approach with 5-fold cross-validation. Feature selection stability was further assessed using 50 iterations of repeated stratified cross-validation. Four machine learning models were constructed: Logistic Regression (LR), Random Forest (RF), XGBoost (XGB), and Support Vector Machine (SVM). Performance was evaluated using ROC curves, PR curves, confusion matrices, and decision curve analysis (DCA). SHAP analysis was applied to the LR model for interpretability. Results A total of 107 radiomics features were extracted, from which 23 key features were selected by LASSO regression. Feature selection stability analysis revealed that 2 features were consistently selected across all 50 iterations (100% frequency), 1 feature was retained in ≥ 98% iterations; overall 6 features exhibited selection frequency ≥ 80%, and 8 features achieved frequency ≥ 50%, while the remaining features showed considerable selection variability. In the internal hold-out test set, the LR model achieved the AUC of 0.926 (95% CI: 0.860–0.978), followed by SVM (AUC = 0.915, 95% CI: 0.848–0.973), RF (AUC = 0.893, 95% CI: 0.814–0.960), and XGBoost (AUC = 0.889, 95% CI: 0.806–0.956). DeLong’s test revealed no statistically significant differences in AUC among the four models (all P > 0.05); the marginal difference between LR and RF (AUC difference = 0.033, P = 0.051) only represented a weak numerical gap without statistical superiority of LR. The LR model demonstrated balanced performance with accuracy 0.836, sensitivity 0.830, specificity 0.850, PPV 0.929, NPV 0.680, F1-score 0.876, and average precision (AP) 0.972; notably, LR yielded the lowest false-negative rate (17.0%), which minimizes the clinical risk of undertreating advanced GC. Hosmer-Lemeshow test revealed good calibration for the Random Forest ( P = 0.586), XGBoost ( P = 0.081), and SVM ( P = 0.213) models, whereas the Logistic Regression model showed evidence of miscalibration ( P < 0.001). Despite this calibration limitation, the LR model’s excellent discriminative performance (AUC = 0.926) supports its clinical utility as a binary classifier. DCA indicated that all four models provided clinical utility, with the LR model offering favorable net benefit across most threshold probabilities. SHAP analysis confirmed that the contribution directions of selected features such as LeastAxisLength and SurfaceVolumeRatio were consistent with pathological interpretations. Conclusion All four machine learning models constructed based on CT venous-phase images showed comparable discriminative power, with no statistically significant inter-model AUC differences detected via DeLong test (all P > 0.05). Rather than possessing statistically superior predictive efficacy, the LR model had higher clinical translation potential due to its excellent interpretability, balanced classification metrics and minimal false-negative risk. The LR model, with transparent coefficient structure and good diagnostic efficacy, can serve as a quantitative auxiliary tool for preoperative staging and treatment decision-making, providing an objective basis for precise GC diagnosis and treatment. Further multi-center external prospective validation is mandatory before formal clinical deployment.

Authors

Publication Details

Journal
Cancer Imaging
Published
2026-10-08
DOI
https://doi.org/10.1186/s40644-026-01132-7
Primary Topic
Radiomics and Machine Learning in Medical Imaging
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

CT venous-phase radiomics for preoperative differentiation of early and advanced gastric cancer: construction, comparison, and validation of machine-learning models

Lei Zhang, Qinpeng Liu, Guowei Zhang, Wei Lian et al.
Cancer Imaging
Radiomics and Machine Learning in Medical Imaging
article

CT venous-phase radiomics for preoperative differentiation of early and advanced gastric cancer: construction, comparison, and validation of machine-learning models

Lei Zhang, Qinpeng Liu, Guowei Zhang, Wei Lian, Ranran Huang, Haijun Zhang, Lin Li, Xinyao Zhao
article en

Abstract

Abstract Objective To construct various radiomics models based on CT venous-phase images for the quantitative assessment of gastric cancer (GC) invasion depth into the gastric wall, and to compare the diagnostic performance of different machine-learning models in distinguishing early (T1–T2 stage) from advanced (T3–T4 stage) GC, aiming to identify a robust and interpretable model and evaluate its potential for clinical application. Methods In this retrospective study, 223 pathologically confirmed GC patients (66 early-stage, 157 advanced-stage) treated between January 2022 and May 2025 were enrolled. Patients were allocated into a training set ( n = 156) and an internal hold-out test set ( n = 67) through stratified random sampling at a 7:3 ratio. All patients underwent enhanced CT within one week prior to surgery. Venous-phase images were selected, and three-dimensional regions of interest (ROIs) encompassing the tumor were manually delineated using 3D Slicer for radiomics feature extraction. To rigorously exclude data leakage, all preprocessing steps—including feature standardization, LASSO feature selection, hyperparameter tuning, and threshold determination—were performed exclusively within the training set using a pipeline-based approach with 5-fold cross-validation. Feature selection stability was further assessed using 50 iterations of repeated stratified cross-validation. Four machine learning models were constructed: Logistic Regression (LR), Random Forest (RF), XGBoost (XGB), and Support Vector Machine (SVM). Performance was evaluated using ROC curves, PR curves, confusion matrices, and decision curve analysis (DCA). SHAP analysis was applied to the LR model for interpretability. Results A total of 107 radiomics features were extracted, from which 23 key features were selected by LASSO regression. Feature selection stability analysis revealed that 2 features were consistently selected across all 50 iterations (100% frequency), 1 feature was retained in ≥ 98% iterations; overall 6 features exhibited selection frequency ≥ 80%, and 8 features achieved frequency ≥ 50%, while the remaining features showed considerable selection variability. In the internal hold-out test set, the LR model achieved the AUC of 0.926 (95% CI: 0.860–0.978), followed by SVM (AUC = 0.915, 95% CI: 0.848–0.973), RF (AUC = 0.893, 95% CI: 0.814–0.960), and XGBoost (AUC = 0.889, 95% CI: 0.806–0.956). DeLong’s test revealed no statistically significant differences in AUC among the four models (all P > 0.05); the marginal difference between LR and RF (AUC difference = 0.033, P = 0.051) only represented a weak numerical gap without statistical superiority of LR. The LR model demonstrated balanced performance with accuracy 0.836, sensitivity 0.830, specificity 0.850, PPV 0.929, NPV 0.680, F1-score 0.876, and average precision (AP) 0.972; notably, LR yielded the lowest false-negative rate (17.0%), which minimizes the clinical risk of undertreating advanced GC. Hosmer-Lemeshow test revealed good calibration for the Random Forest ( P = 0.586), XGBoost ( P = 0.081), and SVM ( P = 0.213) models, whereas the Logistic Regression model showed evidence of miscalibration ( P < 0.001). Despite this calibration limitation, the LR model’s excellent discriminative performance (AUC = 0.926) supports its clinical utility as a binary classifier. DCA indicated that all four models provided clinical utility, with the LR model offering favorable net benefit across most threshold probabilities. SHAP analysis confirmed that the contribution directions of selected features such as LeastAxisLength and SurfaceVolumeRatio were consistent with pathological interpretations. Conclusion All four machine learning models constructed based on CT venous-phase images showed comparable discriminative power, with no statistically significant inter-model AUC differences detected via DeLong test (all P > 0.05). Rather than possessing statistically superior predictive efficacy, the LR model had higher clinical translation potential due to its excellent interpretability, balanced classification metrics and minimal false-negative risk. The LR model, with transparent coefficient structure and good diagnostic efficacy, can serve as a quantitative auxiliary tool for preoperative staging and treatment decision-making, providing an objective basis for precise GC diagnosis and treatment. Further multi-center external prospective validation is mandatory before formal clinical deployment.

Cancer Imaging
Openalex Percentile: Top 13%
Radiomics and Machine Learning in Medical Imaging
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.