The Finnish Diabetes Risk Score (FINDRISC) Combined with Explainable Machine Learning for Detection of Previously Unrecognized Dysglycemia in Primary Care: A CatBoost-Based Approach

Background/Objectives: To evaluate a CatBoost-based machine learning (ML) approach for detecting previously unrecognized dysglycemia at the index primary-care visit and to assess the contribution of the Finnish Diabetes Risk Score (FINDRISC) and routinely available demographic, clinical, and laboratory variables using SHAP. Methods: This retrospective cross-sectional study included 2033 adults without previously diagnosed prediabetes or diabetes. The dataset was randomly split once into a training set of 1626 participants and an internal held-out test set of 407 participants. Three CatBoost feature-set configurations were evaluated: FINDRISC Only, SHAP-selected variables Without FINDRISC, and the same variables With FINDRISC. Five-fold stratified cross-validation and hyperparameter optimization were performed within the training set; the held-out test set was not used for model fitting or tuning. Raw FINDRISC was additionally evaluated as a non-ML benchmark. Test-set evaluation included discrimination, classification metrics, calibration, and paired-bootstrap comparison of AUCs. Results: Overall, 935 (46.0%) participants had screen-detected dysglycemia. On the held-out test set, AUCs were 0.758 for Without FINDRISC, 0.835 for FINDRISC Only, 0.840 for With FINDRISC, and 0.833 for raw FINDRISC. The incremental AUC of With FINDRISC versus FINDRISC Only was 0.005 (95% paired bootstrap CI −0.012 to 0.022; p = 0.556), indicating no statistically significant improvement. Brier scores were 0.203, 0.169, and 0.166 for Without FINDRISC, FINDRISC Only, and With FINDRISC, respectively. Conclusions: FINDRISC accounted for most of the observed discrimination for prevalent, previously unrecognized dysglycemia. Adding routine clinical and laboratory variables within CatBoost produced only a small, statistically non-significant incremental improvement. These findings should be regarded as preliminary internal evidence and require prospective external validation.

Authors

Institutions

Publication Details

Journal
Healthcare
Published
2026-09-29
DOI
https://doi.org/10.3390/healthcare14193214
Primary Topic
Diabetes, Cardiovascular Risks, and Lipoproteins
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

The Finnish Diabetes Risk Score (FINDRISC) Combined with Explainable Machine Learning for Detection of Previously Unrecognized Dysglycemia in Primary Care: A CatBoost-Based Approach

Fevziye Turkoglu Genc, Mehmet Yildiz
Healthcare
Diabetes, Cardiovascular Risks, and Lipoproteins
article

The Finnish Diabetes Risk Score (FINDRISC) Combined with Explainable Machine Learning for Detection of Previously Unrecognized Dysglycemia in Primary Care: A CatBoost-Based Approach

Fevziye Turkoglu Genc, Mehmet Yildiz
article en

Abstract

Background/Objectives: To evaluate a CatBoost-based machine learning (ML) approach for detecting previously unrecognized dysglycemia at the index primary-care visit and to assess the contribution of the Finnish Diabetes Risk Score (FINDRISC) and routinely available demographic, clinical, and laboratory variables using SHAP. Methods: This retrospective cross-sectional study included 2033 adults without previously diagnosed prediabetes or diabetes. The dataset was randomly split once into a training set of 1626 participants and an internal held-out test set of 407 participants. Three CatBoost feature-set configurations were evaluated: FINDRISC Only, SHAP-selected variables Without FINDRISC, and the same variables With FINDRISC. Five-fold stratified cross-validation and hyperparameter optimization were performed within the training set; the held-out test set was not used for model fitting or tuning. Raw FINDRISC was additionally evaluated as a non-ML benchmark. Test-set evaluation included discrimination, classification metrics, calibration, and paired-bootstrap comparison of AUCs. Results: Overall, 935 (46.0%) participants had screen-detected dysglycemia. On the held-out test set, AUCs were 0.758 for Without FINDRISC, 0.835 for FINDRISC Only, 0.840 for With FINDRISC, and 0.833 for raw FINDRISC. The incremental AUC of With FINDRISC versus FINDRISC Only was 0.005 (95% paired bootstrap CI −0.012 to 0.022; p = 0.556), indicating no statistically significant improvement. Brier scores were 0.203, 0.169, and 0.166 for Without FINDRISC, FINDRISC Only, and With FINDRISC, respectively. Conclusions: FINDRISC accounted for most of the observed discrimination for prevalent, previously unrecognized dysglycemia. Adding routine clinical and laboratory variables within CatBoost produced only a small, statistically non-significant incremental improvement. These findings should be regarded as preliminary internal evidence and require prospective external validation.

HealthcareVol. 14(19)
İstanbul Kanuni Sultan Süleyman Eğitim ve Araştırma Hastanesi (TR), Giresun İl Sağlık Müdürlüğü (TR)
Reduced inequalities, Peace, Justice and strong institutions
Openalex Percentile: Top 12%
Diabetes, Cardiovascular Risks, and Lipoproteins
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.