Performance Evaluation of Single and Ensemble Models for Customer Churn Prediction Analysis

Customer churn remains one of the most consequential problems facing business enterprises globally. Compared with the enticement and acquisition costs associated with acquiring new customers, retaining an existing subscriber is substantially cheaper. Due to its significance to business sustainability, various studies have analysed churn using various predictive analytics approaches. However, the majority of the proposed models rarely consider the key contribution of class imbalance issues, present controlled comparisons across single and ensemble models, or provide explanations and interpretation contexts to the predictions. In this study, nine classification algorithms, including six single learners (Logistic Regression (LR), Decision Tree (DT), Gaussian Naive Bayes (Gaussian NB), K-Nearest Neighbours (K-NN), Support Vector Machine (SVM), Random Forest (RF)) and three boosting ensembles (Gradient Boosting Machine (GBM), Extreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM)), are evaluated under an identical preprocessing and validation protocol using the publicly available churn dataset of 7043 subscribers, with repeated stratified fivefold cross-validation and five random restarts. Class imbalance is addressed through Adaptive Synthetic Sampling (ADASYN), with its contribution compared against no resampling, class weighting, and Synthetic Minority Over-sampling Technique (SMOTE). The best-performing configuration, LightGBM under ADASYN, achieved a mean F1 of 0.7212, an ROC-AUC of 0.9094, and a PR-AUC of 0.7916, outperforming other models by a statistically detectable but practically modest margin. The LightGBM is further subjected to a two-level interpretability analysis using SHapley Additive exPlanations (SHAP) for global reasoning and Local Interpretable Model-agnostic Explanations (LIME) for case-level reasoning. The resulting insights are translated into a retention framework that links churn drivers to practical business decisions.

Authors

Institutions

Publication Details

Journal
Informatics
Published
2026-09-30
DOI
https://doi.org/10.3390/informatics13100159
Primary Topic
Customer churn and segmentation
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Performance Evaluation of Single and Ensemble Models for Customer Churn Prediction Analysis

Oludayo Olufolorunsho Olugbara, Oyeniyi Akeem Alimi, Smangele Pretty Moyane, Fatima Labake Ajani
Informatics
Customer churn and segmentation
article

Performance Evaluation of Single and Ensemble Models for Customer Churn Prediction Analysis

Oludayo Olufolorunsho Olugbara, Oyeniyi Akeem Alimi, Smangele Pretty Moyane, Fatima Labake Ajani
article en

Abstract

Customer churn remains one of the most consequential problems facing business enterprises globally. Compared with the enticement and acquisition costs associated with acquiring new customers, retaining an existing subscriber is substantially cheaper. Due to its significance to business sustainability, various studies have analysed churn using various predictive analytics approaches. However, the majority of the proposed models rarely consider the key contribution of class imbalance issues, present controlled comparisons across single and ensemble models, or provide explanations and interpretation contexts to the predictions. In this study, nine classification algorithms, including six single learners (Logistic Regression (LR), Decision Tree (DT), Gaussian Naive Bayes (Gaussian NB), K-Nearest Neighbours (K-NN), Support Vector Machine (SVM), Random Forest (RF)) and three boosting ensembles (Gradient Boosting Machine (GBM), Extreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM)), are evaluated under an identical preprocessing and validation protocol using the publicly available churn dataset of 7043 subscribers, with repeated stratified fivefold cross-validation and five random restarts. Class imbalance is addressed through Adaptive Synthetic Sampling (ADASYN), with its contribution compared against no resampling, class weighting, and Synthetic Minority Over-sampling Technique (SMOTE). The best-performing configuration, LightGBM under ADASYN, achieved a mean F1 of 0.7212, an ROC-AUC of 0.9094, and a PR-AUC of 0.7916, outperforming other models by a statistically detectable but practically modest margin. The LightGBM is further subjected to a two-level interpretability analysis using SHapley Additive exPlanations (SHAP) for global reasoning and Local Interpretable Model-agnostic Explanations (LIME) for case-level reasoning. The resulting insights are translated into a retention framework that links churn drivers to practical business decisions.

InformaticsVol. 13(10)
Durban University of Technology (ZA)
Responsible consumption and production
Openalex Percentile: Top 6%
Customer churn and segmentation
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.