Prediction versus inference: A dual-pipeline analytical framework for evaluating GAIN-based imputation in hypertension risk modeling

Generative Adversarial Imputation Networks (GAIN) are increasingly used to address missing data and improve clinical prediction models, but whether GAIN-imputed data preserve reliable regression-based inference remains unclear. This study compared two analytical strategies for hypertension risk modeling: a complete-case analytical dataset (n = 19,860) and a GAIN-imputed analytical dataset (n = 31,072). LASSO, generalized linear modeling (GLM), and XGBoost were applied to evaluate predictive performance and regression-based inference. Model performance was assessed using AUC, calibration curves, and decision curve analysis, while interpretability and stability were evaluated using adjusted odds ratios, SHAP values, univariate AUC, Bland–Altman analysis, and bootstrap resampling. Compared with the complete-case analysis, the GAIN-imputed analytical dataset yielded higher predictive performance across all models, with the largest improvement observed for XGBoost (AUC: 0.752 to 0.801), together with improved calibration and clinical utility. However, differences were also observed in inference-oriented analyses. Several biomarkers demonstrated attenuation, inflation, or reversal of adjusted odds ratio direction between the two analytical strategies. Potassium shifted from protective to risk-associated, while albumin and total calcium also showed directional reversal. Magnesium further illustrated the limitation of linear inference, with SHAP dependence analysis suggesting a nonlinear U-shaped relationship despite a positive GLM coefficient. These findings highlight the distinction between prediction-oriented and inference-oriented analyses when applying GAIN-based imputation in electronic health record data. Rather than demonstrating a direct causal effect of imputation on regression-based inference, our results show that analytical conclusions may differ substantially between complete-case and GAIN-imputed datasets. We therefore propose a dual-pipeline framework in which GAIN-imputed data support prediction-oriented modeling and automated hypertension risk scoring, whereas the complete-case analytical dataset remains the primary reference for regression-based inference and biomarker interpretation. This framework provides a practical strategy for integrating generative AI into EHR-based clinical decision support while preserving transparency and reproducibility.

Authors

Institutions

Publication Details

Journal
Informatics in Medicine Unlocked
Published
2026-09-21
DOI
https://doi.org/10.1016/j.imu.2026.101818
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Prediction versus inference: A dual-pipeline analytical framework for evaluating GAIN-based imputation in hypertension risk modeling

Gong Zhang, Han Yong-sheng, Weiwei Liao, Jun Fu et al.
Informatics in Medicine Unlocked
Artificial Intelligence in Healthcare and Education
article

Prediction versus inference: A dual-pipeline analytical framework for evaluating GAIN-based imputation in hypertension risk modeling

Gong Zhang, Han Yong-sheng, Weiwei Liao, Jun Fu, Alan Chan, Fei Liu, Tao Song
article en

Abstract

Generative Adversarial Imputation Networks (GAIN) are increasingly used to address missing data and improve clinical prediction models, but whether GAIN-imputed data preserve reliable regression-based inference remains unclear. This study compared two analytical strategies for hypertension risk modeling: a complete-case analytical dataset (n = 19,860) and a GAIN-imputed analytical dataset (n = 31,072). LASSO, generalized linear modeling (GLM), and XGBoost were applied to evaluate predictive performance and regression-based inference. Model performance was assessed using AUC, calibration curves, and decision curve analysis, while interpretability and stability were evaluated using adjusted odds ratios, SHAP values, univariate AUC, Bland–Altman analysis, and bootstrap resampling. Compared with the complete-case analysis, the GAIN-imputed analytical dataset yielded higher predictive performance across all models, with the largest improvement observed for XGBoost (AUC: 0.752 to 0.801), together with improved calibration and clinical utility. However, differences were also observed in inference-oriented analyses. Several biomarkers demonstrated attenuation, inflation, or reversal of adjusted odds ratio direction between the two analytical strategies. Potassium shifted from protective to risk-associated, while albumin and total calcium also showed directional reversal. Magnesium further illustrated the limitation of linear inference, with SHAP dependence analysis suggesting a nonlinear U-shaped relationship despite a positive GLM coefficient. These findings highlight the distinction between prediction-oriented and inference-oriented analyses when applying GAIN-based imputation in electronic health record data. Rather than demonstrating a direct causal effect of imputation on regression-based inference, our results show that analytical conclusions may differ substantially between complete-case and GAIN-imputed datasets. We therefore propose a dual-pipeline framework in which GAIN-imputed data support prediction-oriented modeling and automated hypertension risk scoring, whereas the complete-case analytical dataset remains the primary reference for regression-based inference and biomarker interpretation. This framework provides a practical strategy for integrating generative AI into EHR-based clinical decision support while preserving transparency and reproducibility.

Informatics in Medicine UnlockedVol. 66
Monash University Malaysia (MY), University of Science and Technology of China (CN), Anhui University (CN), Harbin Medical University (CN), First Affiliated Hospital of Harbin Medical University (CN), The University of Winnipeg (CA)
Peace, Justice and strong institutions
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.