SHAP-guided adaptive feature weighting for phishing URL detection with a federated learning extension
Abstract Machine-learning phishing URL detectors routinely report accuracies above 97%, yet four weaknesses undermine these results: absent statistical testing, unfair latency benchmarking, uninterpretable predictions, and an inability to use data held privately across organisations. We address all four on the Hannousse-Yahiouche benchmark (11,430 URLs, 87 features). We introduce the SHAP-Guided Adaptive Feature Weighting (SAFW) layer, a trainable component initialised from XGBoost SHAP scores and evaluate seven architectures (including a DistilBERT reference case) with 10-fold cross-validation, twenty-one McNemar tests, five class-imbalance ratios, and CPU-only latency benchmarking. XGBoost dominates centralised training (97.00% accuracy, 0.79 ms latency); no statistically detectable difference was observed between SVM and BiLSTM (McNemar p = 0.469). Extending to a twenty-client federated setting (matching the N ≥ 20 convention in the federated-learning literature) under non-IID Dirichlet partitioning with a per-client sample guard, federated SAFW-Hybrid reaches 95.4% ± 0.5% over the final rounds on the same 80/20 split, showing no statistically detectable difference from its centralised counterpart (95.45%; McNemar p = 0.30; paired-difference 95% CI [− 1.05, + 0.22] pp against a ± 0.5 pp margin not formally equivalent, as the CI is not wholly contained within the margin). A naive federated XGBoost ensemble loses 4.2 points, but a purpose-built horizontal federated GBDT (cyclic federated XGBoost) recovers almost all of this (94.5–95.1%), matching federated SAFW-Hybrid within test-set noise. Excluding the two externally-sourced features (google_index, page_rank) leaves centralised and federated accuracy largely intact (within ~ 1% point) but sharply reduces class-imbalance robustness, with F1 at the most extreme ratio variying by 35.6% points depending on architecture. Cross-client analysis shows google_index remains highly stable across the twenty simulated clients under the evaluated non-IID FedAvg setting (mean weight 0.891, spread 0.0013) under a non-saturating initialisation, ruling out sigmoid saturation as the cause; because the clients are coupled through repeated FedAvg aggregation rather than trained in isolation, this stability is reported as a property of the evaluated setting rather than independent cross-client convergence. XGBoost therefore remains the accuracy leader in the centralised setting; no statistically detectable difference was observed between XGBoost and SAFW-Hybrid under the evaluated federated setting once aggregation is done properly; the federated value of SAFW-Hybrid is competitive accuracy combined with an interpretable, cross-client-stable per-feature weight vector that a tree ensemble cannot provide.
Authors
- Mohammad Aknan (ORCID: https://orcid.org/0000-0003-1608-5574)
- Priyanshu priyanshu (ORCID: https://orcid.org/0009-0004-3201-4666)
- Madiha Saleem
- Kahkashan Kouser
- Onkar Singh
Institutions
- Central Forensic Science Laboratory (IN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1038/s41598-026-71965-6
- Primary Topic
- Spam and Phishing Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00