Explainable Ensemble Learning for Phishing URL Detection: A Comparative and Interpretability-Driven Evaluation
Phishing attacks continue to grow in scale and sophistication, using deceptive URLs to imitate legitimate platforms and compromise sensitive user data.Blacklist-and rule-based defenses may struggle to detect newly emerging and short-lived phishing domains, creating a need for data-driven detection approaches.This paper presents a comparative and explainable machine learning framework for phishing URL detection using lexical and domain-level URL features.Four machine learning classifiers, namely Decision Tree, Random Forest, Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost), are evaluated using publicly available phishing URL datasets.The models are compared using accuracy, precision, recall, F1-score, and Receiver Operating Characteristic Area Under the Curve (ROC-AUC).To improve model transparency, SHapley Additive exPlanations (SHAP) are incorporated to quantify the contribution of individual URL features to model predictions.The study further analyzes the most influential features associated with phishing classification and compares the interpretability of the selected models.The proposed approach aims to combine effective phishing detection with transparent and understandable predictions, enabling cybersecurity practitioners to examine the factors influencing classification decisions.The findings are expected to support the development of practical, auditable, and reproducible phishing URL detection systems.
Authors
- Vikrant Satish Salunkhe
- Pratiksha Rajendra Dashpute
- Sheetal Shrikant Shevkari
Institutions
- MIT Art, Design and Technology University (IN)
Publication Details
- Journal
- International Journal of Innovative Research in Technology
- Published
- 2026-09-16
- DOI
- https://doi.org/10.64643/ijirt.208542-459
- Primary Topic
- Spam and Phishing Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00