Cluster-aware imputation: a hybrid XGBoost-based model for missing data
Abstract The fast growth of high-dimensional datasets in computational biology, industrial IoT, and socio-economic research has made the issue of missing data more challenging. This common problem reduces statistical power and introduces systematic bias in later machine learning tasks. Traditional methods, such as listwise deletion and univariate mean substitution, fail to keep the multivariate covariance structures that are required for accurate decisions. In contrast to traditional Multivariate Imputation by Chained Equations (MICE) frameworks whereby only global regression models that are based on the assumption of single data distribution are mostly used, the modern MICE framework provides improvements through inter-variable modeling. This assumption is not very accurate when it comes to hidden subpopulations and local variances in real datasets. The present work introduces a novel approach in the form of the Cluster Imputation Framework (ClusteringImputer) that is a hybrid model which incorporates unsupervised dimensionality reduction and density-based clustering with supervised gradient boosting. This method specifically captures the local correlation structures by first partitioning the feature space into similar subgroups using Principal Component Analysis and K-Means clustering and then training cluster-specific XGBoost regressors. We accompany a thorough benchmarking analysis of 97 datasets from the UCI Machine Learning Repository; thus, our framework is proven to achieve asymptotic linearity of $$O\left(N\right)$$ in computational complexity. This stands in strong opposition to the K-Nearest Neighbor's (KNN) quadratic $$O\left({N}^{2}\right)$$ scaling. The findings show that despite the fact that instance-based methods work well on continuous low-dimensional manifolds, the Cluster-Aware approach is more efficient with better reconstruction accuracy and scalability in the large scale with more than $$\left(N>{10}^{4}\right)$$ instances, particularly in mixed-type and categorical-dominant situations.
Authors
- Mai Mohamed (ORCID: https://orcid.org/0009-0000-4488-8089)
- Mohamed Emad
- Mahmoud M. Ismail
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1038/s41598-026-70170-9
- Primary Topic
- Statistical Methods and Bayesian Inference
- Type
- article
- Field-Weighted Citation Impact
- 0.00