Comparative Analysis of Machine Learning Algorithms for Phishing Website Detection
Phishing remains one of the most prevalent and financially damaging cyber threats, in which attackers impersonate legitimate websites to steal sensitive user information such as passwords and banking credentials. This paper investigates the use of supervised machine learning algorithms to automatically detect phishing websites using the UCI Phishing Websites dataset, which contains 11,055 instances described by 30 URL- and page-based features. Three classification algorithms (Logistic Regression, Random Forest, and Support Vector Machine) were trained and evaluated on an 80/20 random train-test split and further validated with 5-fold cross-validation. Random Forest achieved the highest test accuracy at 96.70%, outperforming SVM (94.71%) and Logistic Regression (92.45%). A feature importance analysis revealed that the final SSL certificate state and the proportion of anchor links pointing away from the domain were the most influential predictors of phishing behavior. Code and data: https://github.com/eln2mac-has/phishing-detection-ml
Authors
- Elnazeer Dawod Sulman
Institutions
- University of Tripoli (LY)
- Center for Solar Energy Research and Studies (LY)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23173758
- Primary Topic
- Spam and Phishing Detection
- Type
- preprint