SHAP-guided adaptive feature weighting for phishing URL detection with a federated learning extension

Abstract Machine-learning phishing URL detectors routinely report accuracies above 97%, yet four weaknesses undermine these results: absent statistical testing, unfair latency benchmarking, uninterpretable predictions, and an inability to use data held privately across organisations. We address all four on the Hannousse-Yahiouche benchmark (11,430 URLs, 87 features). We introduce the SHAP-Guided Adaptive Feature Weighting (SAFW) layer, a trainable component initialised from XGBoost SHAP scores and evaluate seven architectures (including a DistilBERT reference case) with 10-fold cross-validation, twenty-one McNemar tests, five class-imbalance ratios, and CPU-only latency benchmarking. XGBoost dominates centralised training (97.00% accuracy, 0.79 ms latency); no statistically detectable difference was observed between SVM and BiLSTM (McNemar p = 0.469). Extending to a twenty-client federated setting (matching the N ≥ 20 convention in the federated-learning literature) under non-IID Dirichlet partitioning with a per-client sample guard, federated SAFW-Hybrid reaches 95.4% ± 0.5% over the final rounds on the same 80/20 split, showing no statistically detectable difference from its centralised counterpart (95.45%; McNemar p = 0.30; paired-difference 95% CI [− 1.05, + 0.22] pp against a ± 0.5 pp margin not formally equivalent, as the CI is not wholly contained within the margin). A naive federated XGBoost ensemble loses 4.2 points, but a purpose-built horizontal federated GBDT (cyclic federated XGBoost) recovers almost all of this (94.5–95.1%), matching federated SAFW-Hybrid within test-set noise. Excluding the two externally-sourced features (google_index, page_rank) leaves centralised and federated accuracy largely intact (within ~ 1% point) but sharply reduces class-imbalance robustness, with F1 at the most extreme ratio variying by 35.6% points depending on architecture. Cross-client analysis shows google_index remains highly stable across the twenty simulated clients under the evaluated non-IID FedAvg setting (mean weight 0.891, spread 0.0013) under a non-saturating initialisation, ruling out sigmoid saturation as the cause; because the clients are coupled through repeated FedAvg aggregation rather than trained in isolation, this stability is reported as a property of the evaluated setting rather than independent cross-client convergence. XGBoost therefore remains the accuracy leader in the centralised setting; no statistically detectable difference was observed between XGBoost and SAFW-Hybrid under the evaluated federated setting once aggregation is done properly; the federated value of SAFW-Hybrid is competitive accuracy combined with an interpretable, cross-client-stable per-feature weight vector that a tree ensemble cannot provide.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-10-05
DOI
https://doi.org/10.1038/s41598-026-71965-6
Primary Topic
Spam and Phishing Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

SHAP-guided adaptive feature weighting for phishing URL detection with a federated learning extension

Mohammad Aknan, Priyanshu priyanshu, Madiha Saleem, Kahkashan Kouser et al.
Scientific Reports
Spam and Phishing Detection
article

SHAP-guided adaptive feature weighting for phishing URL detection with a federated learning extension

Mohammad Aknan, Priyanshu priyanshu, Madiha Saleem, Kahkashan Kouser, Onkar Singh
article en

Abstract

Abstract Machine-learning phishing URL detectors routinely report accuracies above 97%, yet four weaknesses undermine these results: absent statistical testing, unfair latency benchmarking, uninterpretable predictions, and an inability to use data held privately across organisations. We address all four on the Hannousse-Yahiouche benchmark (11,430 URLs, 87 features). We introduce the SHAP-Guided Adaptive Feature Weighting (SAFW) layer, a trainable component initialised from XGBoost SHAP scores and evaluate seven architectures (including a DistilBERT reference case) with 10-fold cross-validation, twenty-one McNemar tests, five class-imbalance ratios, and CPU-only latency benchmarking. XGBoost dominates centralised training (97.00% accuracy, 0.79 ms latency); no statistically detectable difference was observed between SVM and BiLSTM (McNemar p = 0.469). Extending to a twenty-client federated setting (matching the N ≥ 20 convention in the federated-learning literature) under non-IID Dirichlet partitioning with a per-client sample guard, federated SAFW-Hybrid reaches 95.4% ± 0.5% over the final rounds on the same 80/20 split, showing no statistically detectable difference from its centralised counterpart (95.45%; McNemar p = 0.30; paired-difference 95% CI [− 1.05, + 0.22] pp against a ± 0.5 pp margin not formally equivalent, as the CI is not wholly contained within the margin). A naive federated XGBoost ensemble loses 4.2 points, but a purpose-built horizontal federated GBDT (cyclic federated XGBoost) recovers almost all of this (94.5–95.1%), matching federated SAFW-Hybrid within test-set noise. Excluding the two externally-sourced features (google_index, page_rank) leaves centralised and federated accuracy largely intact (within ~ 1% point) but sharply reduces class-imbalance robustness, with F1 at the most extreme ratio variying by 35.6% points depending on architecture. Cross-client analysis shows google_index remains highly stable across the twenty simulated clients under the evaluated non-IID FedAvg setting (mean weight 0.891, spread 0.0013) under a non-saturating initialisation, ruling out sigmoid saturation as the cause; because the clients are coupled through repeated FedAvg aggregation rather than trained in isolation, this stability is reported as a property of the evaluated setting rather than independent cross-client convergence. XGBoost therefore remains the accuracy leader in the centralised setting; no statistically detectable difference was observed between XGBoost and SAFW-Hybrid under the evaluated federated setting once aggregation is done properly; the federated value of SAFW-Hybrid is competitive accuracy combined with an interpretable, cross-client-stable per-feature weight vector that a tree ensemble cannot provide.

Scientific Reports
Central Forensic Science Laboratory (IN)
Openalex Percentile: Top 8%
Spam and Phishing Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.