Cluster-aware imputation: a hybrid XGBoost-based model for missing data

Abstract The fast growth of high-dimensional datasets in computational biology, industrial IoT, and socio-economic research has made the issue of missing data more challenging. This common problem reduces statistical power and introduces systematic bias in later machine learning tasks. Traditional methods, such as listwise deletion and univariate mean substitution, fail to keep the multivariate covariance structures that are required for accurate decisions. In contrast to traditional Multivariate Imputation by Chained Equations (MICE) frameworks whereby only global regression models that are based on the assumption of single data distribution are mostly used, the modern MICE framework provides improvements through inter-variable modeling. This assumption is not very accurate when it comes to hidden subpopulations and local variances in real datasets. The present work introduces a novel approach in the form of the Cluster Imputation Framework (ClusteringImputer) that is a hybrid model which incorporates unsupervised dimensionality reduction and density-based clustering with supervised gradient boosting. This method specifically captures the local correlation structures by first partitioning the feature space into similar subgroups using Principal Component Analysis and K-Means clustering and then training cluster-specific XGBoost regressors. We accompany a thorough benchmarking analysis of 97 datasets from the UCI Machine Learning Repository; thus, our framework is proven to achieve asymptotic linearity of $$O\left(N\right)$$ in computational complexity. This stands in strong opposition to the K-Nearest Neighbor's (KNN) quadratic $$O\left({N}^{2}\right)$$ scaling. The findings show that despite the fact that instance-based methods work well on continuous low-dimensional manifolds, the Cluster-Aware approach is more efficient with better reconstruction accuracy and scalability in the large scale with more than $$\left(N>{10}^{4}\right)$$ instances, particularly in mixed-type and categorical-dominant situations.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-10-07
DOI
https://doi.org/10.1038/s41598-026-70170-9
Primary Topic
Statistical Methods and Bayesian Inference
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Cluster-aware imputation: a hybrid XGBoost-based model for missing data

Mai Mohamed, Mohamed Emad, Mahmoud M. Ismail
Scientific Reports
Statistical Methods and Bayesian Inference
article

Cluster-aware imputation: a hybrid XGBoost-based model for missing data

Mai Mohamed, Mohamed Emad, Mahmoud M. Ismail
article en

Abstract

Abstract The fast growth of high-dimensional datasets in computational biology, industrial IoT, and socio-economic research has made the issue of missing data more challenging. This common problem reduces statistical power and introduces systematic bias in later machine learning tasks. Traditional methods, such as listwise deletion and univariate mean substitution, fail to keep the multivariate covariance structures that are required for accurate decisions. In contrast to traditional Multivariate Imputation by Chained Equations (MICE) frameworks whereby only global regression models that are based on the assumption of single data distribution are mostly used, the modern MICE framework provides improvements through inter-variable modeling. This assumption is not very accurate when it comes to hidden subpopulations and local variances in real datasets. The present work introduces a novel approach in the form of the Cluster Imputation Framework (ClusteringImputer) that is a hybrid model which incorporates unsupervised dimensionality reduction and density-based clustering with supervised gradient boosting. This method specifically captures the local correlation structures by first partitioning the feature space into similar subgroups using Principal Component Analysis and K-Means clustering and then training cluster-specific XGBoost regressors. We accompany a thorough benchmarking analysis of 97 datasets from the UCI Machine Learning Repository; thus, our framework is proven to achieve asymptotic linearity of $$O\left(N\right)$$ in computational complexity. This stands in strong opposition to the K-Nearest Neighbor's (KNN) quadratic $$O\left({N}^{2}\right)$$ scaling. The findings show that despite the fact that instance-based methods work well on continuous low-dimensional manifolds, the Cluster-Aware approach is more efficient with better reconstruction accuracy and scalability in the large scale with more than $$\left(N>{10}^{4}\right)$$ instances, particularly in mixed-type and categorical-dominant situations.

Scientific ReportsVol. 16(1)
Openalex Percentile: Top 10%
Statistical Methods and Bayesian Inference
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Cluster-aware imputation: a hybrid XGBoost-based model for missing data — Mai Mohamed, Mohamed Emad, et al. · Scientific Reports (2026) | TGRS Research Map | TGRS