Explainable machine learning on repeated screening data for risk-stratified diabetes care: a dual-level interpretability approach

In South Korea, national health screening intervals are determined by occupational category: typically, annual screening for non-office workers and biennial screening for office workers. To enable proactive prevention, it is critical to understand how these screening intervals and an individual’s baseline diabetes status influence disease progression. This study aimed to identify distinct clinical patterns within this regulatory framework and to establish a pre-emptive intervention strategy using machine learning (ML). From an initial pool of 73,767 participants (133,387 records) collected between 2014 and 2020, 25,275 individuals were included in the final analysis following rigorous preprocessing and exclusion. Participants were stratified into four subsets based on their baseline diabetes status and screening intervals (1-year vs. 2-year). Six ML algorithms were compared, and Shapley Additive exPlanations (SHAP) analysis was used to provide multi-level interpretability, ranging from population-level feature importance to individual-level risk factors. Tree-based ensemble models demonstrated favorable predictive performance, with clinical differences more pronounced in the 2-year interval than in the 1-year interval. Piecewise regression analysis of SHAP dependence plots successfully identified population-level model-derived breakpoints (thresholds) for major predictors, reflecting interval-associated metabolic differences. Furthermore, individual-level risk assessments using SHAP force plots highlighted the model’s potential for personalized clinical counseling. Our findings emphasize that aligning monitoring strategies with an individual’s initial diabetes status and considering screening interval may be useful for proactive prevention. By providing both population-level model-derived breakpoints and individualized risk profiles, this study offers an explainable model-derived framework to support early intervention before the onset of overt diabetes.

Authors

Institutions

Publication Details

Journal
BMC Medical Informatics and Decision Making
Published
2026-09-17
DOI
https://doi.org/10.1186/s12911-026-03847-w
Primary Topic
Diabetes, Cardiovascular Risks, and Lipoproteins
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Explainable machine learning on repeated screening data for risk-stratified diabetes care: a dual-level interpretability approach

Tae‐Young Heo, Jung Kee Min, Hyeonseop Yuk, Wonmi Gu et al.
BMC Medical Informatics and Decision Making
Diabetes, Cardiovascular Risks, and Lipoproteins
article

Explainable machine learning on repeated screening data for risk-stratified diabetes care: a dual-level interpretability approach

Tae‐Young Heo, Jung Kee Min, Hyeonseop Yuk, Wonmi Gu, Jaesuk Yun
article en

Abstract

In South Korea, national health screening intervals are determined by occupational category: typically, annual screening for non-office workers and biennial screening for office workers. To enable proactive prevention, it is critical to understand how these screening intervals and an individual’s baseline diabetes status influence disease progression. This study aimed to identify distinct clinical patterns within this regulatory framework and to establish a pre-emptive intervention strategy using machine learning (ML). From an initial pool of 73,767 participants (133,387 records) collected between 2014 and 2020, 25,275 individuals were included in the final analysis following rigorous preprocessing and exclusion. Participants were stratified into four subsets based on their baseline diabetes status and screening intervals (1-year vs. 2-year). Six ML algorithms were compared, and Shapley Additive exPlanations (SHAP) analysis was used to provide multi-level interpretability, ranging from population-level feature importance to individual-level risk factors. Tree-based ensemble models demonstrated favorable predictive performance, with clinical differences more pronounced in the 2-year interval than in the 1-year interval. Piecewise regression analysis of SHAP dependence plots successfully identified population-level model-derived breakpoints (thresholds) for major predictors, reflecting interval-associated metabolic differences. Furthermore, individual-level risk assessments using SHAP force plots highlighted the model’s potential for personalized clinical counseling. Our findings emphasize that aligning monitoring strategies with an individual’s initial diabetes status and considering screening interval may be useful for proactive prevention. By providing both population-level model-derived breakpoints and individualized risk profiles, this study offers an explainable model-derived framework to support early intervention before the onset of overt diabetes.

BMC Medical Informatics and Decision Making
Ulsan College (KR), Chungbuk National University (KR), University of Ulsan (KR), Ulsan University Hospital (KR)
Reduced inequalities
Openalex Percentile: Top 11%
Diabetes, Cardiovascular Risks, and Lipoproteins
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.