A secure and explainable phishing detection framework using deep learning and LLM-based reasoning
Existing phishing detection systems face a persistent trade-off: feature-based classical models offer interpretability but sacrifice accuracy on evolving attack patterns, while deep learning models achieve strong accuracy but function as opaque black boxes that provide security analysts with only a label and a confidence score, not actionable reasoning. Recent efforts to close this gap with large language models (LLMs) introduce their own risks hallucination and prompt injection that remain largely unaddressed in prior phishing-detection literature. To resolve these limitations, this paper proposes a security-aware and explainable phishing detection framework combining deep learning, explainable artificial intelligence (XAI), and evidence-grounded LLM reasoning. The proposed framework fuses raw URL sequence representations (via a CNN–BiLSTM network) with twenty handcrafted lexical and domain-based statistical indicators through a multi-branch deep architecture. To provide actionable transparency, LIME and SHAP are applied to a Random Forest surrogate trained on the handcrafted feature space to generate local and global feature-level attributions, complementing the deep hybrid model’s predictions with verifiable, model-derived evidence. These explanations are subsequently structured as standardized evidence packets to drive an LLM reasoning module that produces analyst-style incident reports with high semantic faithfulness (0.833) and robust risk-severity consistency ( $$100\%$$ ). To ensure operational reliability, the framework incorporates a security-hardened LLM inference pipeline featuring regex input sanitization, structured prompting, and strict programmatic JSON schema validation; red-team evaluations across twelve attack vectors confirm these mechanisms meaningfully reduce susceptibility to adversarial prompt injection ( $$0\%$$ schema violation). Experimental evaluation on a primary benchmark ( $$n=428{,}616$$ URLs) and an independent cross-dataset test set ( $$n=142{,}859$$ URLs) demonstrates that the proposed framework achieves an F1-score of 0.9926 and a ROC-AUC of 0.9999, matching or exceeding both classical and deep learning baselines while exhibiting superior generalization under distribution shift.
Authors
- Hoc Minh Le (ORCID: https://orcid.org/0009-0003-7986-6922)
- Van Thuong Nguyen
Institutions
- Eskisehir Technical University (TR)
- Konya Technical University (TR)
Publication Details
- Journal
- Discover Informatics
- Published
- 2026-10-01
- DOI
- https://doi.org/10.1007/s44564-026-00020-3
- Primary Topic
- Spam and Phishing Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00