Association of the CIMI with multimorbidity classification among U.S. adults with CKD and cancer via interpretable SHAP-based machine learning
Abstract People who carry a diagnosis of both chronic kidney disease (CKD) and cancer face a markedly greater likelihood of unfavorable clinical outcomes, a situation that complicates therapeutic planning and everyday disease management. A pathological process common to the two disorders is sustained inflammation, which drives disease advancement through disturbed immune regulation, oxidative injury and altered metabolism. Composite indices that pool inflammatory and metabolic signals from routine hematological and biochemical tests could therefore help to recognize, among individuals already diagnosed with either CKD or cancer, those who harbor both conditions at the same time. Nonetheless, classification studies undertaken with the specific purpose of separating CKD–cancer comorbidity from either condition alone remain scarce. Drawing on data from the National Health and Nutrition Examination Survey (NHANES), we set out to appraise systematically how well a panel of Comprehensive Inflammatory and Metabolic Indices (CIMI) discriminates concurrent CKD–cancer comorbidity, and to construct and internally assess an optimal classifier using several machine learning algorithms. Records for 139,205 adults were first examined across several NHANES waves spanning 1999 through August 2023. Adopting the comorbidity-oriented selection approaches applied in earlier NHANES investigations, the analysis finally comprised 1,385 adults affected by CKD together with cancer and 9,506 adults affected by only one of the two conditions (10,891 adults in all following complete-case screening), every one of whom possessed complete laboratory and diagnostic information. Thirty-one CIMI—computed from peripheral blood cell counts, serum albumin, liver function parameters and lipid profiles—were merged with demographic variables. Classification models were then built with six algorithms: random forest, Extreme Gradient Boosting (XGBoost), Logistic Regression (LR), Recursive Partitioning and Regression Trees (RPART), Naive Bayes (NB) and k-nearest neighbors (KNN). Development and evaluation proceeded under stratified 10-fold cross-validation applied to the whole analytic sample ( n = 10,891). SMOTE was confined to the interior of each training fold, and every performance statistic was derived from pooled out-of-fold predictions. Neither a holdout set nor an external validation cohort was employed, and the primary endpoint was the area under the receiver operating characteristic curve (AUC-ROC). Contributions of individual features within the best model were interpreted using SHapley Additive exPlanations (SHAP). This work addressed a cross-sectional classification of a pre-existing composite comorbidity state rather than forecasting future disease. Since CKD status rested on one examination only, the outcome denotes concurrent cancer accompanied by diminished kidney function and/or albuminuria, not CKD confirmed clinically; likewise, because cancer status came exclusively from a self-reported questionnaire item (MCQ220) that carried no information on subtype, stage or treatment, the outcome constitutes a heterogeneous composite label. XGBoost delivered the strongest overall performance of the six algorithms, attaining an AUC-ROC of 0.776 (95% CI 0.763–0.789), an AUPRC of 0.445 (95% CI 0.415–0.471) and a Matthews Correlation Coefficient (MCC) of 0.287; every paired DeLong comparison with the remaining five algorithms proved significant (all P < 0.001). Sensitivity reached 0.665 and specificity 0.735 at the Youden-optimal threshold (0.133), while at the conventional 0.5 cut-off the corresponding values were 0.227 and 0.987. According to the SHAP analysis, the ten most influential features comprised Age, CALLY, CLR, monocyte count (MON), FIB-4, ALB, Race_3, SIRI/AISI, IBI and CRP; each one remained within the top ten in ≥ 93% of 100 resampling replicates, yet, since a number of these indices are mathematically nested, such attributions are exploratory and cannot be construed as independent clinical contributions. With the NHANES database, we constructed and internally assessed a classification model for CKD–cancer comorbidity that combines several CIMI. XGBoost ranked first in discriminative ability among the six algorithms examined. Sensitivity nevertheless stayed poor at the conventional 0.5 threshold (0.227), and probability calibration was only moderate (Brier score 0.092; calibration slope 0.79); the model thus cannot serve on its own as a screening or diagnostic instrument and is better viewed as an exploratory classifier of a pre-existing composite state. Age, CALLY and CLR also emerged as important discriminative features. These results shed new light on the classification of people living with CKD and cancer concurrently, on the behavior of classification models and the contribution of explainable modeling to comorbidity research, and on the design of more targeted clinical assessment approaches.
Authors
- 潘金伸
- Xubo Dai (ORCID: https://orcid.org/0009-0003-2436-6066)
- ChaoLan Fang
- WeiXiong Xu (ORCID: https://orcid.org/0009-0002-3429-4767)
- YongJian Ye
- PiBo Du
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1038/s41598-026-73936-3
- Primary Topic
- Inflammatory Biomarkers in Disease Prognosis
- Type
- article
- Field-Weighted Citation Impact
- 0.00