Interpretable Machine Learning for Preliminary Lung Cancer Risk Assessment Based on Clinical Indicators and Lifestyle Factors

Background/Objective: Lung cancer remains a leading cause of cancer-related mortality, and access to low-dose CT screening is limited in resource-constrained settings. This study aimed to develop and evaluate an interpretable machine learning model for preliminary, symptom-based lung cancer risk discrimination using clinical indicators and lifestyle factors, without requiring imaging. Methods: A cross-sectional dataset of 561 participants (38 with a prior lung cancer diagnosis), following exclusion of a minors-eligible age bracket for which independent age verification was not possible, was analyzed. Five machine learning algorithms (Logistic Regression, Random Forest, XGBoost, LightGBM, CatBoost) were compared using a leakage-free pipeline, with ADASYN class balancing and hyperparameter tuning performed exclusively within cross-validation folds; feature selection was verified to be stable across folds (98% average overlap with the final feature set). The final model was selected solely on the basis of repeated cross-validated performance, without reference to the test set. Model interpretability was assessed using LIME, and a four-tier risk stratification system was developed and calibrated. Results: Random Forest achieved the best cross-validated F1-score (0.850 ± 0.092) among the five candidates and, given a clinically motivated preference for higher sensitivity at comparable F1, was selected as the final model. On the independent test set (n = 113), it achieved Accuracy = 0.956, Recall = 0.875, F1-score = 0.737, and AUC-ROC = 0.912. Sensitivity analysis indicated moderate reliance on a potential diagnostic proxy feature, and balancing-method comparisons (ADASYN, random oversampling, class weighting) confirmed that conclusions were not contingent on the specific resampling technique. Conclusions: The proposed model demonstrates solid discriminative performance and interpretability; however, given the cross-sectional design and limited number of positive cases, results should be interpreted as preliminary risk discrimination rather than validated early-detection prediction, warranting further prospective validation.

Authors

Institutions

Publication Details

Journal
Computers
Published
2026-10-09
DOI
https://doi.org/10.3390/computers15100695
Primary Topic
Artificial Intelligence in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Interpretable Machine Learning for Preliminary Lung Cancer Risk Assessment Based on Clinical Indicators and Lifestyle Factors

Dinara Kozhakhmetova, Alina Bugubayeva, Dinara Shyrynkhanova, Lazzat Kydyralina et al.
Computers
Artificial Intelligence in Healthcare
article

Interpretable Machine Learning for Preliminary Lung Cancer Risk Assessment Based on Clinical Indicators and Lifestyle Factors

Dinara Kozhakhmetova, Alina Bugubayeva, Dinara Shyrynkhanova, Lazzat Kydyralina, Shynggys Adilgazyuly, Indira Karymsakova, Dariga Bekenova
article en

Abstract

Background/Objective: Lung cancer remains a leading cause of cancer-related mortality, and access to low-dose CT screening is limited in resource-constrained settings. This study aimed to develop and evaluate an interpretable machine learning model for preliminary, symptom-based lung cancer risk discrimination using clinical indicators and lifestyle factors, without requiring imaging. Methods: A cross-sectional dataset of 561 participants (38 with a prior lung cancer diagnosis), following exclusion of a minors-eligible age bracket for which independent age verification was not possible, was analyzed. Five machine learning algorithms (Logistic Regression, Random Forest, XGBoost, LightGBM, CatBoost) were compared using a leakage-free pipeline, with ADASYN class balancing and hyperparameter tuning performed exclusively within cross-validation folds; feature selection was verified to be stable across folds (98% average overlap with the final feature set). The final model was selected solely on the basis of repeated cross-validated performance, without reference to the test set. Model interpretability was assessed using LIME, and a four-tier risk stratification system was developed and calibrated. Results: Random Forest achieved the best cross-validated F1-score (0.850 ± 0.092) among the five candidates and, given a clinically motivated preference for higher sensitivity at comparable F1, was selected as the final model. On the independent test set (n = 113), it achieved Accuracy = 0.956, Recall = 0.875, F1-score = 0.737, and AUC-ROC = 0.912. Sensitivity analysis indicated moderate reliance on a potential diagnostic proxy feature, and balancing-method comparisons (ADASYN, random oversampling, class weighting) confirmed that conclusions were not contingent on the specific resampling technique. Conclusions: The proposed model demonstrates solid discriminative performance and interpretability; however, given the cross-sectional design and limited number of positive cases, results should be interpreted as preliminary risk discrimination rather than validated early-detection prediction, warranting further prospective validation.

ComputersVol. 15(10)
D. Serikbayev East Kazakhstan State Technical University (KZ), Private Institution “Karaganda University of Kazpotrebsoyuz” (KZ), Semey Medical University (KZ), Astana Medical University (KZ), Shakarim University (KZ)
Openalex Percentile: Top 8%
Artificial Intelligence in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.