SHAP-based interpretation of machine learning models for cardiometabolic risk in fatty liver disease
Abstract Fatty liver disease has become the most common chronic liver condition worldwide, yet the dominant threat to affected patients is cardiovascular rather than hepatic. Whether their excess cardiovascular risk tracks the amount of hepatic fat or the cardiometabolic and fibrotic burden that accompanies it remains unresolved–an ambiguity that conventional risk scores, derived in the general population, are poorly equipped to address. We examined this question with an interpretable machine-learning analysis of two complementary outcomes, studied in two distinct fatty-liver cohorts drawn from a nationally representative health survey with linked mortality records and using only routinely collected clinical and laboratory variables, including self-reported lipid-lowering and antihypertensive therapy. The first task was cross-sectional classification of prevalent cardiovascular disease ( $$n=3{,}217$$ n = 3 , 217 adults with elastography-defined steatosis; 414 prevalent cases, $$12.9\\%$$ 12.9 % ); the second was longitudinal prediction of all-cause mortality ( $$n=16{,}178$$ n = 16 , 178 adults with Fatty Liver Index–defined steatosis; 1,679 deaths over a median 7.2 years, of which 535 were cardiovascular). Seven classifiers and four survival learners were trained over a shared feature space and preprocessing protocol; gradient-boosted trees, fitted on the natural event prevalence, were retained as the reference model for each task and explained through a common TreeSHAP attribution layer. Discrimination was strong (held-out test AUROC up to 0.837, reference model 0.828, with high negative predictive value, well-calibrated probabilities, and net benefit on decision-curve analysis); the survival model reached a concordance index of 0.852 and separated predicted-risk tertiles into markedly distinct 15-year survival (approximately $$98\\%$$ 98 % , $$92\\%$$ 92 % , and $$43\\%$$ 43 % ). Crucially, the controlled attenuation parameter, liver stiffness, and the Fatty Liver Index did not separate the outcome groups and received little attribution weight, whereas predictions were dominated by age, hepatic fibrosis (FIB-4), smoking, renal function, treatment status, and social deprivation. The full models outperformed parsimonious clinical predictors and remained robust across steatosis definitions and clinical subgroups. Within these models, the degree of steatosis contributed little to prediction relative to the cardiometabolic burden that accompanies it–an association, not a causal verdict–offering a transparent, inexpensive basis for risk stratification from routine data that now warrants external and prospective validation.
Authors
- Yong Li (ORCID: https://orcid.org/0000-0002-7230-3196)
- Lu Li
- Gang Zhao
- Haitao Shi
- Longbao Yang
- Wenni Zhuang
Institutions
- Second Affiliated Hospital of Xi'an Jiaotong University (CN)
Publication Details
- Journal
- Journal of King Saud University - Computer and Information Sciences
- Published
- 2026-08-26
- DOI
- https://doi.org/10.1007/s44443-026-01179-3
- Primary Topic
- Liver Disease Diagnosis and Treatment
- Type
- article
- Field-Weighted Citation Impact
- 0.00