Reproducible benchmarking of four machine-learning classifiers on a large synthetic heart-disease competition dataset

Background Public synthetic datasets enable transparent model comparison but cannot establish performance in clinical populations. We asked how four fixed classifiers compare under one leakage-controlled validation framework on a large synthetic heart-disease competition dataset. Methods The official training set contained 630,000 observations (44.83% outcome prevalence) and 13 predictors; the 270,000-row test set had no outcome labels. L2-regularized logistic regression, HistGradientBoosting, CatBoost, and LightGBM were evaluated using identical stratified five-fold outer splits. Logistic-regression preprocessing was fitted within outer-training data, and early stopping used outer-training data only. The primary measure was pooled out-of-fold (OOF) AUROC; secondary measures were average precision, Brier score, calibration intercept and slope, fold variability, and train–validation optimism. Duplicate/overlap, identifier-only, and permuted-outcome controls assessed leakage. Results Pooled OOF AUROC was 0.955443 for CatBoost, 0.955324 for LightGBM, 0.955066 for HistGradientBoosting, and 0.952874 for logistic regression. Corresponding average precision ranged from 0.945912 to 0.948828 and Brier scores from 0.081116 to 0.083518. Calibration intercepts ranged from −0.001299 to 0.002153 and slopes from 0.998896 to 1.004388. Mean optimism was 0.000017–0.002686. No duplicated identifiers, duplicated feature profiles, or train–test feature-profile overlap were found; identifier-only AUROC was 0.500121 and mean permuted-outcome AUROC was 0.500615. Conclusions All four models discriminated strongly and showed close internal calibration, while gradient-boosting models had only marginally higher AUROC than logistic regression. The principal contribution is a fully rerunnable, controlled benchmark rather than a novel algorithm or a clinically validated prediction model. External performance in real clinical cohorts remains unknown.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-21
DOI
https://doi.org/10.1371/journal.pone.0354304
Primary Topic
Advanced Causal Inference Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Reproducible benchmarking of four machine-learning classifiers on a large synthetic heart-disease competition dataset

Raed M. Ennab
PLoS ONE
Advanced Causal Inference Techniques
article

Reproducible benchmarking of four machine-learning classifiers on a large synthetic heart-disease competition dataset

Raed M. Ennab
article en

Abstract

Background Public synthetic datasets enable transparent model comparison but cannot establish performance in clinical populations. We asked how four fixed classifiers compare under one leakage-controlled validation framework on a large synthetic heart-disease competition dataset. Methods The official training set contained 630,000 observations (44.83% outcome prevalence) and 13 predictors; the 270,000-row test set had no outcome labels. L2-regularized logistic regression, HistGradientBoosting, CatBoost, and LightGBM were evaluated using identical stratified five-fold outer splits. Logistic-regression preprocessing was fitted within outer-training data, and early stopping used outer-training data only. The primary measure was pooled out-of-fold (OOF) AUROC; secondary measures were average precision, Brier score, calibration intercept and slope, fold variability, and train–validation optimism. Duplicate/overlap, identifier-only, and permuted-outcome controls assessed leakage. Results Pooled OOF AUROC was 0.955443 for CatBoost, 0.955324 for LightGBM, 0.955066 for HistGradientBoosting, and 0.952874 for logistic regression. Corresponding average precision ranged from 0.945912 to 0.948828 and Brier scores from 0.081116 to 0.083518. Calibration intercepts ranged from −0.001299 to 0.002153 and slopes from 0.998896 to 1.004388. Mean optimism was 0.000017–0.002686. No duplicated identifiers, duplicated feature profiles, or train–test feature-profile overlap were found; identifier-only AUROC was 0.500121 and mean permuted-outcome AUROC was 0.500615. Conclusions All four models discriminated strongly and showed close internal calibration, while gradient-boosting models had only marginally higher AUROC than logistic regression. The principal contribution is a fully rerunnable, controlled benchmark rather than a novel algorithm or a clinically validated prediction model. External performance in real clinical cohorts remains unknown.

PLoS ONEVol. 21(9)
Yarmouk University (JO)
Reduced inequalities
Openalex Percentile: Top 8%
Advanced Causal Inference Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Reproducible benchmarking of four machine-learning classifiers on a large synthetic heart-disease competition dataset — Raed M. Ennab · PLoS ONE (2026) | TGRS Research Map | TGRS