The Explainability–Reliability Gap in Fraud Detection: Evidence from SHAP and Permutation Importance Under Distribution Shift

This study develops an empirical audit framework for assessing the explainability–reliability gap in fraud detection: whether stable model explanations remain consistent with performance-based feature reliance under distribution shift. Using the Bank Account Fraud dataset suite, including a Base dataset and five biased variants, the study examines group-size disparity, fraud-prevalence disparity, separability bias, and temporal shift. Logistic Regression, Linear SVC, and Random Forest are benchmarked using standard classification metrics, followed by cross-variant evaluation with Logistic Regression as the interpretable baseline. SHAP is used to assess explanation stability, while permutation importance measures performance-based feature reliance. The results show that accuracy and ROC-AUC can overstate practical effectiveness under severe class imbalance; notably, Random Forest retained useful discrimination while producing near-zero recall at the evaluated threshold. SHAP feature rankings remained relatively stable across variants, particularly for address-history, identity-similarity, credit-risk, device, and behavioral variables. However, permutation importance revealed weaker and more variable reliance on several SHAP-ranked features. The limited agreement between the two measures indicates a partial explainability–reliability gap. The findings show that explanation stability alone is insufficient for evaluating trustworthy fraud detection models and should be complemented by performance-based validation under biased and shifted deployment conditions.

Authors

Institutions

Publication Details

Journal
FinTech
Published
2026-09-06
DOI
https://doi.org/10.3390/fintech5030077
Primary Topic
Imbalanced Data Classification Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

The Explainability–Reliability Gap in Fraud Detection: Evidence from SHAP and Permutation Importance Under Distribution Shift

Tanvir Bhuiyan, Ariful Hoque, Rahma Mirza, Istiaque Bhuiyan
FinTech
Imbalanced Data Classification Techniques
article

The Explainability–Reliability Gap in Fraud Detection: Evidence from SHAP and Permutation Importance Under Distribution Shift

Tanvir Bhuiyan, Ariful Hoque, Rahma Mirza, Istiaque Bhuiyan
article en

Abstract

This study develops an empirical audit framework for assessing the explainability–reliability gap in fraud detection: whether stable model explanations remain consistent with performance-based feature reliance under distribution shift. Using the Bank Account Fraud dataset suite, including a Base dataset and five biased variants, the study examines group-size disparity, fraud-prevalence disparity, separability bias, and temporal shift. Logistic Regression, Linear SVC, and Random Forest are benchmarked using standard classification metrics, followed by cross-variant evaluation with Logistic Regression as the interpretable baseline. SHAP is used to assess explanation stability, while permutation importance measures performance-based feature reliance. The results show that accuracy and ROC-AUC can overstate practical effectiveness under severe class imbalance; notably, Random Forest retained useful discrimination while producing near-zero recall at the evaluated threshold. SHAP feature rankings remained relatively stable across variants, particularly for address-history, identity-similarity, credit-risk, device, and behavioral variables. However, permutation importance revealed weaker and more variable reliance on several SHAP-ranked features. The limited agreement between the two measures indicates a partial explainability–reliability gap. The findings show that explanation stability alone is insufficient for evaluating trustworthy fraud detection models and should be complemented by performance-based validation under biased and shifted deployment conditions.

FinTechVol. 5(3)
Murdoch University (AU), Bangladesh University (BD), La Trobe University (AU)
Peace, Justice and strong institutions, Reduced inequalities
Openalex Percentile: Top 8%
Imbalanced Data Classification Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

The Explainability–Reliability Gap in Fraud Detection: Evidence from SHAP and Permutation Importance Under Distribution Shift — Tanvir Bhuiyan, Ariful Hoque, et al. · FinTech (2026) | TGRS Research Map | TGRS