Prediction versus inference: A dual-pipeline analytical framework for evaluating GAIN-based imputation in hypertension risk modeling
Generative Adversarial Imputation Networks (GAIN) are increasingly used to address missing data and improve clinical prediction models, but whether GAIN-imputed data preserve reliable regression-based inference remains unclear. This study compared two analytical strategies for hypertension risk modeling: a complete-case analytical dataset (n = 19,860) and a GAIN-imputed analytical dataset (n = 31,072). LASSO, generalized linear modeling (GLM), and XGBoost were applied to evaluate predictive performance and regression-based inference. Model performance was assessed using AUC, calibration curves, and decision curve analysis, while interpretability and stability were evaluated using adjusted odds ratios, SHAP values, univariate AUC, Bland–Altman analysis, and bootstrap resampling. Compared with the complete-case analysis, the GAIN-imputed analytical dataset yielded higher predictive performance across all models, with the largest improvement observed for XGBoost (AUC: 0.752 to 0.801), together with improved calibration and clinical utility. However, differences were also observed in inference-oriented analyses. Several biomarkers demonstrated attenuation, inflation, or reversal of adjusted odds ratio direction between the two analytical strategies. Potassium shifted from protective to risk-associated, while albumin and total calcium also showed directional reversal. Magnesium further illustrated the limitation of linear inference, with SHAP dependence analysis suggesting a nonlinear U-shaped relationship despite a positive GLM coefficient. These findings highlight the distinction between prediction-oriented and inference-oriented analyses when applying GAIN-based imputation in electronic health record data. Rather than demonstrating a direct causal effect of imputation on regression-based inference, our results show that analytical conclusions may differ substantially between complete-case and GAIN-imputed datasets. We therefore propose a dual-pipeline framework in which GAIN-imputed data support prediction-oriented modeling and automated hypertension risk scoring, whereas the complete-case analytical dataset remains the primary reference for regression-based inference and biomarker interpretation. This framework provides a practical strategy for integrating generative AI into EHR-based clinical decision support while preserving transparency and reproducibility.
Authors
- Gong Zhang (ORCID: https://orcid.org/0000-0002-3651-3639)
- Han Yong-sheng
- Weiwei Liao
- Jun Fu (ORCID: https://orcid.org/0009-0007-3976-6408)
- Alan Chan (ORCID: https://orcid.org/0009-0003-4649-2805)
- Fei Liu
- Tao Song (ORCID: https://orcid.org/0000-0002-7501-2001)
Institutions
- Monash University Malaysia (MY)
- University of Science and Technology of China (CN)
- Anhui University (CN)
- Harbin Medical University (CN)
- First Affiliated Hospital of Harbin Medical University (CN)
- The University of Winnipeg (CA)
Publication Details
- Journal
- Informatics in Medicine Unlocked
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1016/j.imu.2026.101818
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00