SHAP-based interpretation of machine learning models for cardiometabolic risk in fatty liver disease

Abstract Fatty liver disease has become the most common chronic liver condition worldwide, yet the dominant threat to affected patients is cardiovascular rather than hepatic. Whether their excess cardiovascular risk tracks the amount of hepatic fat or the cardiometabolic and fibrotic burden that accompanies it remains unresolved–an ambiguity that conventional risk scores, derived in the general population, are poorly equipped to address. We examined this question with an interpretable machine-learning analysis of two complementary outcomes, studied in two distinct fatty-liver cohorts drawn from a nationally representative health survey with linked mortality records and using only routinely collected clinical and laboratory variables, including self-reported lipid-lowering and antihypertensive therapy. The first task was cross-sectional classification of prevalent cardiovascular disease ( $$n=3{,}217$$ n = 3 , 217 adults with elastography-defined steatosis; 414 prevalent cases, $$12.9\\%$$ 12.9 % ); the second was longitudinal prediction of all-cause mortality ( $$n=16{,}178$$ n = 16 , 178 adults with Fatty Liver Index–defined steatosis; 1,679 deaths over a median 7.2 years, of which 535 were cardiovascular). Seven classifiers and four survival learners were trained over a shared feature space and preprocessing protocol; gradient-boosted trees, fitted on the natural event prevalence, were retained as the reference model for each task and explained through a common TreeSHAP attribution layer. Discrimination was strong (held-out test AUROC up to 0.837, reference model 0.828, with high negative predictive value, well-calibrated probabilities, and net benefit on decision-curve analysis); the survival model reached a concordance index of 0.852 and separated predicted-risk tertiles into markedly distinct 15-year survival (approximately $$98\\%$$ 98 % , $$92\\%$$ 92 % , and $$43\\%$$ 43 % ). Crucially, the controlled attenuation parameter, liver stiffness, and the Fatty Liver Index did not separate the outcome groups and received little attribution weight, whereas predictions were dominated by age, hepatic fibrosis (FIB-4), smoking, renal function, treatment status, and social deprivation. The full models outperformed parsimonious clinical predictors and remained robust across steatosis definitions and clinical subgroups. Within these models, the degree of steatosis contributed little to prediction relative to the cardiometabolic burden that accompanies it–an association, not a causal verdict–offering a transparent, inexpensive basis for risk stratification from routine data that now warrants external and prospective validation.

Authors

Institutions

Publication Details

Journal
Journal of King Saud University - Computer and Information Sciences
Published
2026-08-26
DOI
https://doi.org/10.1007/s44443-026-01179-3
Primary Topic
Liver Disease Diagnosis and Treatment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SHAP-based interpretation of machine learning models for cardiometabolic risk in fatty liver disease

Yong Li, Lu Li, Gang Zhao, Haitao Shi et al.
Journal of King Saud University - Computer and Information Sciences
Liver Disease Diagnosis and Treatment
article

SHAP-based interpretation of machine learning models for cardiometabolic risk in fatty liver disease

Yong Li, Lu Li, Gang Zhao, Haitao Shi, Longbao Yang, Wenni Zhuang
article en

Abstract

Abstract Fatty liver disease has become the most common chronic liver condition worldwide, yet the dominant threat to affected patients is cardiovascular rather than hepatic. Whether their excess cardiovascular risk tracks the amount of hepatic fat or the cardiometabolic and fibrotic burden that accompanies it remains unresolved–an ambiguity that conventional risk scores, derived in the general population, are poorly equipped to address. We examined this question with an interpretable machine-learning analysis of two complementary outcomes, studied in two distinct fatty-liver cohorts drawn from a nationally representative health survey with linked mortality records and using only routinely collected clinical and laboratory variables, including self-reported lipid-lowering and antihypertensive therapy. The first task was cross-sectional classification of prevalent cardiovascular disease ( $$n=3{,}217$$ n = 3 , 217 adults with elastography-defined steatosis; 414 prevalent cases, $$12.9\%$$ 12.9 % ); the second was longitudinal prediction of all-cause mortality ( $$n=16{,}178$$ n = 16 , 178 adults with Fatty Liver Index–defined steatosis; 1,679 deaths over a median 7.2 years, of which 535 were cardiovascular). Seven classifiers and four survival learners were trained over a shared feature space and preprocessing protocol; gradient-boosted trees, fitted on the natural event prevalence, were retained as the reference model for each task and explained through a common TreeSHAP attribution layer. Discrimination was strong (held-out test AUROC up to 0.837, reference model 0.828, with high negative predictive value, well-calibrated probabilities, and net benefit on decision-curve analysis); the survival model reached a concordance index of 0.852 and separated predicted-risk tertiles into markedly distinct 15-year survival (approximately $$98\%$$ 98 % , $$92\%$$ 92 % , and $$43\%$$ 43 % ). Crucially, the controlled attenuation parameter, liver stiffness, and the Fatty Liver Index did not separate the outcome groups and received little attribution weight, whereas predictions were dominated by age, hepatic fibrosis (FIB-4), smoking, renal function, treatment status, and social deprivation. The full models outperformed parsimonious clinical predictors and remained robust across steatosis definitions and clinical subgroups. Within these models, the degree of steatosis contributed little to prediction relative to the cardiometabolic burden that accompanies it–an association, not a causal verdict–offering a transparent, inexpensive basis for risk stratification from routine data that now warrants external and prospective validation.

Journal of King Saud University - Computer and Information SciencesVol. 38(7)
Second Affiliated Hospital of Xi'an Jiaotong University (CN)
Peace, Justice and strong institutions, Reduced inequalities
Openalex Percentile: Top 10%
Liver Disease Diagnosis and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.