Loan default risk classification in digital banking services: evaluating autoencoder-based feature learning and gradient boosting

Abstract Loan default remains a persistent source of financial loss in digital banking, and prior work in this area has reported near-perfect classification results that rarely hold up under close scrutiny. This study re-examines one such case. Using the publicly available “Loan Status Classification” dataset on Kaggle (100,000 loan records; https://www.kaggle.com/code/sazack/loan-status-classification ), we first show that a commonly used derived label in this dataset, a five-category credit status rating, is a deterministic bucketing of the credit score feature itself (reconstructable at 100% accuracy from that single feature), making it unsuitable as a classification target since the “nswer” is directly encoded in an input variable. We instead use the dataset's genuine outcome label, loan repayment status (Fully Paid vs. Charged Off), and evaluate eight classification approaches—Decision Tree, Random Forest, a Multilayer Perceptron, a Naive Bayes classifier, Gradient Boosting, Gradient Boosting after Principal Component Analysis, Gradient Boosting after Autoencoder-based feature learning, and a majority vote Ensemble—under a leakage-controlled, deduplicated, stratified 5-fold cross-validation protocol in which every preprocessing step (imputation, class rebalancing, and dimensionality reduction) is fit exclusively on each training fold. Under this corrected protocol, no model approaches perfect performance, and the Autoencoder-based hybrid does not outperform simpler alternatives: plain Gradient Boosting achieves the strongest overall balance of metrics (F1 = 0.495, ROC-AUC = 0.764, PR-AUC = 0.595), with Random Forest and a simple Ensemble close behind, while the Autoencoder-plus-Gradient-Boosting hybrid trails most baselines (F1 = 0.348, ROC-AUC = 0.621)—differences confirmed as statistically significant by paired McNemar tests on the pooled out-of-fold predictions. These findings offer a concrete, verified illustration of how an unexamined derived label can manufacture apparent model superiority, and provide a corrected, reproducible benchmark for loan default risk classification together with practical guidance for deploying credit risk models responsibly in digital banking services.

Authors

Institutions

Publication Details

Journal
Future Business Journal
Published
2026-09-25
DOI
https://doi.org/10.1186/s43093-026-01011-4
Primary Topic
Financial Distress and Bankruptcy Prediction
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Loan default risk classification in digital banking services: evaluating autoencoder-based feature learning and gradient boosting

Fahad Alsheref
Future Business Journal
Financial Distress and Bankruptcy Prediction
article

Loan default risk classification in digital banking services: evaluating autoencoder-based feature learning and gradient boosting

Fahad Alsheref
article en

Abstract

Abstract Loan default remains a persistent source of financial loss in digital banking, and prior work in this area has reported near-perfect classification results that rarely hold up under close scrutiny. This study re-examines one such case. Using the publicly available “Loan Status Classification” dataset on Kaggle (100,000 loan records; https://www.kaggle.com/code/sazack/loan-status-classification ), we first show that a commonly used derived label in this dataset, a five-category credit status rating, is a deterministic bucketing of the credit score feature itself (reconstructable at 100% accuracy from that single feature), making it unsuitable as a classification target since the “nswer” is directly encoded in an input variable. We instead use the dataset's genuine outcome label, loan repayment status (Fully Paid vs. Charged Off), and evaluate eight classification approaches—Decision Tree, Random Forest, a Multilayer Perceptron, a Naive Bayes classifier, Gradient Boosting, Gradient Boosting after Principal Component Analysis, Gradient Boosting after Autoencoder-based feature learning, and a majority vote Ensemble—under a leakage-controlled, deduplicated, stratified 5-fold cross-validation protocol in which every preprocessing step (imputation, class rebalancing, and dimensionality reduction) is fit exclusively on each training fold. Under this corrected protocol, no model approaches perfect performance, and the Autoencoder-based hybrid does not outperform simpler alternatives: plain Gradient Boosting achieves the strongest overall balance of metrics (F1 = 0.495, ROC-AUC = 0.764, PR-AUC = 0.595), with Random Forest and a simple Ensemble close behind, while the Autoencoder-plus-Gradient-Boosting hybrid trails most baselines (F1 = 0.348, ROC-AUC = 0.621)—differences confirmed as statistically significant by paired McNemar tests on the pooled out-of-fold predictions. These findings offer a concrete, verified illustration of how an unexamined derived label can manufacture apparent model superiority, and provide a corrected, reproducible benchmark for loan default risk classification together with practical guidance for deploying credit risk models responsibly in digital banking services.

Future Business JournalVol. 12(1)
King Khalid University (SA)
Openalex Percentile: Top 4%
Financial Distress and Bankruptcy Prediction
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.