The Explainability–Reliability Gap in Fraud Detection: Evidence from SHAP and Permutation Importance Under Distribution Shift
This study develops an empirical audit framework for assessing the explainability–reliability gap in fraud detection: whether stable model explanations remain consistent with performance-based feature reliance under distribution shift. Using the Bank Account Fraud dataset suite, including a Base dataset and five biased variants, the study examines group-size disparity, fraud-prevalence disparity, separability bias, and temporal shift. Logistic Regression, Linear SVC, and Random Forest are benchmarked using standard classification metrics, followed by cross-variant evaluation with Logistic Regression as the interpretable baseline. SHAP is used to assess explanation stability, while permutation importance measures performance-based feature reliance. The results show that accuracy and ROC-AUC can overstate practical effectiveness under severe class imbalance; notably, Random Forest retained useful discrimination while producing near-zero recall at the evaluated threshold. SHAP feature rankings remained relatively stable across variants, particularly for address-history, identity-similarity, credit-risk, device, and behavioral variables. However, permutation importance revealed weaker and more variable reliance on several SHAP-ranked features. The limited agreement between the two measures indicates a partial explainability–reliability gap. The findings show that explanation stability alone is insufficient for evaluating trustworthy fraud detection models and should be complemented by performance-based validation under biased and shifted deployment conditions.
Authors
- Tanvir Bhuiyan (ORCID: https://orcid.org/0000-0002-3379-9731)
- Ariful Hoque (ORCID: https://orcid.org/0000-0001-8369-6653)
- Rahma Mirza
- Istiaque Bhuiyan (ORCID: https://orcid.org/0009-0009-1992-5445)
Institutions
- Murdoch University (AU)
- Bangladesh University (BD)
- La Trobe University (AU)
Publication Details
- Journal
- FinTech
- Published
- 2026-09-06
- DOI
- https://doi.org/10.3390/fintech5030077
- Primary Topic
- Imbalanced Data Classification Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00