Explainable machine learning models for AI enabled credit scoring and financial inclusion in emerging markets
This study examines whether explainable machine-learning approaches can support more inclusive credit scoring while maintaining acceptable predictive performance, transparency, and fairness in an emerging-market microfinance setting. We analyzed a de-identified dataset of 2500 resolved loan applications obtained from participating microfinance institutions in the MENA region during 2022–2023. Nine supervised algorithms were compared, spanning classical benchmarks (Logistic Regression, Decision Tree, Random Forest, Gradient Boosting, Support Vector Machine, K-Nearest Neighbors) and modern gradient-boosting frameworks (XGBoost, LightGBM, CatBoost). Default (90 + days past due) was modeled as the positive class. Performance was evaluated on a held-out test set using ROC-AUC and PR-AUC with bootstrap confidence intervals, supported by repeated stratified cross-validation and DeLong tests. Interpretability was assessed with SHAP, including stability and cross-model consistency checks, and fairness was evaluated using disparate impact, equal opportunity, predictive parity, group-level calibration, and intersectional analysis. An ablation design isolated the incremental contribution of non-traditional indicators. CatBoost achieved the highest test-set discrimination (ROC-AUC = 0.783, 95% CI [0.741, 0.825]; PR-AUC = 0.506) but was statistically indistinguishable from regularized Logistic Regression (ROC-AUC = 0.780; DeLong p = .59), and no gradient-boosting model significantly outperformed the logistic benchmark. SHAP analysis identified Credit Score Category, Utility Payment Behavior, and Repayment History as the dominant predictors, with rankings highly stable under resampling (Spearman ρ = 0.999) and consistent across model families (ρ = 0.95–0.98 among tree ensembles). Adding non-traditional indicators produced a positive but statistically non-significant incremental gain (ΔAUC = + 0.008, 95% CI [− 0.010, 0.027]). Disparate impact ratios exceeded the 0.80 heuristic for Gender (0.956), Geographic Region (0.961), and all gender-by-region intersections (0.881), with small equal-opportunity and predictive-parity gaps, although rural applicants experienced a higher false-positive rate than urban applicants (0.680 vs. 0.492). In this microfinance context, transparent and well-calibrated models matched the discrimination of complex ensembles, non-traditional indicators carried weak evidence of incremental predictive value despite meaningful within-model signal but limited incremental discrimination, and headline parity metrics coexisted with subgroup error-rate asymmetries. These findings support explainability- and governance-centered model selection for inclusive credit scoring, while underscoring the need for multi-metric fairness auditing, external validation, and continued monitoring in operational deployment.
Authors
- Anas Ahmad Bani Atta (ORCID: https://orcid.org/0000-0003-1064-3117)
- Yazan Taher Shawabkeh (ORCID: https://orcid.org/0009-0005-4468-869X)
- Ahmad Marei (ORCID: https://orcid.org/0000-0002-7976-3013)
- Atef Badri Al-Qur'an
- Rana Alofishat
- Khaled Abidallah Aldarabah
Institutions
- National Agricultural Research Center (JO)
- Middle East University (JO)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1007/s44163-026-02392-9
- Primary Topic
- Microfinance and Financial Inclusion
- Type
- article
- Field-Weighted Citation Impact
- 0.00