Interpretable Machine Learning for Preliminary Lung Cancer Risk Assessment Based on Clinical Indicators and Lifestyle Factors
Background/Objective: Lung cancer remains a leading cause of cancer-related mortality, and access to low-dose CT screening is limited in resource-constrained settings. This study aimed to develop and evaluate an interpretable machine learning model for preliminary, symptom-based lung cancer risk discrimination using clinical indicators and lifestyle factors, without requiring imaging. Methods: A cross-sectional dataset of 561 participants (38 with a prior lung cancer diagnosis), following exclusion of a minors-eligible age bracket for which independent age verification was not possible, was analyzed. Five machine learning algorithms (Logistic Regression, Random Forest, XGBoost, LightGBM, CatBoost) were compared using a leakage-free pipeline, with ADASYN class balancing and hyperparameter tuning performed exclusively within cross-validation folds; feature selection was verified to be stable across folds (98% average overlap with the final feature set). The final model was selected solely on the basis of repeated cross-validated performance, without reference to the test set. Model interpretability was assessed using LIME, and a four-tier risk stratification system was developed and calibrated. Results: Random Forest achieved the best cross-validated F1-score (0.850 ± 0.092) among the five candidates and, given a clinically motivated preference for higher sensitivity at comparable F1, was selected as the final model. On the independent test set (n = 113), it achieved Accuracy = 0.956, Recall = 0.875, F1-score = 0.737, and AUC-ROC = 0.912. Sensitivity analysis indicated moderate reliance on a potential diagnostic proxy feature, and balancing-method comparisons (ADASYN, random oversampling, class weighting) confirmed that conclusions were not contingent on the specific resampling technique. Conclusions: The proposed model demonstrates solid discriminative performance and interpretability; however, given the cross-sectional design and limited number of positive cases, results should be interpreted as preliminary risk discrimination rather than validated early-detection prediction, warranting further prospective validation.
Authors
- Dinara Kozhakhmetova (ORCID: https://orcid.org/0000-0002-4327-3899)
- Alina Bugubayeva
- Dinara Shyrynkhanova
- Lazzat Kydyralina
- Shynggys Adilgazyuly
- Indira Karymsakova
- Dariga Bekenova
Institutions
- D. Serikbayev East Kazakhstan State Technical University (KZ)
- Private Institution “Karaganda University of Kazpotrebsoyuz” (KZ)
- Semey Medical University (KZ)
- Astana Medical University (KZ)
- Shakarim University (KZ)
Publication Details
- Journal
- Computers
- Published
- 2026-10-09
- DOI
- https://doi.org/10.3390/computers15100695
- Primary Topic
- Artificial Intelligence in Healthcare
- Type
- article
- Field-Weighted Citation Impact
- 0.00