ShieldPhish: Evaluating Representation Robustness and Cross-Source Generalization in Phishing URL Detection

Phishing URL detectors frequently report strong performance under conventional offline evaluations, yet such estimates maynot transfer to independently sourced operational traffic. In this study, we evaluate the representation robustness and cross-sourcegeneralization of lexical phishing classifiers. We diagnose a severe representation sensitivity in conventional models, where trivialformatting shortcuts (e.g., the presence of https://) artificially inflate offline accuracy while causing near 100% predictioninstability across URL variants. To mitigate this, we introduce a component-aware canonicalization pipeline and a hybrid rule fusion framework, completely eliminating representation-based class flips and reducing the false positive rate (FPR) on secondarybenchmarks to 0.6%. However, to evaluate true operational readiness, we enforced a strict, hash-verified, version-controlled precollection freeze protocol utilizing domain-disjoint post-freeze external snapshots (OpenPhish and a post-freeze Cisco UmbrellaTop 1M snapshot). Despite our successful representation repair and strong secondary benchmarks, the frozen model exhibited amassive cross-source distribution shift, maintaining 91.34% phishing recall (95% CI: 87.4–94.8%) but degrading to an 86.84%FPR (95% CI: 83.6–90.0%) on independently sourced Umbrella benign observations. The combined ROC-AUC was 0.460 (95%CI: 0.392–0.524), providing no evidence of useful ranking discrimination under this source composition. Post hoc error analysisrevealed that, while complex subdomains and infrastructure reverse-DNS records were heavily enriched failure modes, the primarydriver was a broad vocabulary and score distribution shift across the unseen benign source. Our findings rigorously demonstrate thateliminating representation shortcuts is necessary, but insufficient, for cross-source generalization, highlighting a critical limitationin relying on constrained offline accuracy for operational deployment assessments.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-04
DOI
https://doi.org/10.5281/zenodo.23141304
Primary Topic
Spam and Phishing Detection
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

ShieldPhish: Evaluating Representation Robustness and Cross-Source Generalization in Phishing URL Detection

Gandhi Raj Giri
Zenodo (CERN European Organization for Nuclear Research)
Spam and Phishing Detection
preprint

ShieldPhish: Evaluating Representation Robustness and Cross-Source Generalization in Phishing URL Detection

Gandhi Raj Giri
preprint en

Abstract

Phishing URL detectors frequently report strong performance under conventional offline evaluations, yet such estimates maynot transfer to independently sourced operational traffic. In this study, we evaluate the representation robustness and cross-sourcegeneralization of lexical phishing classifiers. We diagnose a severe representation sensitivity in conventional models, where trivialformatting shortcuts (e.g., the presence of https://) artificially inflate offline accuracy while causing near 100% predictioninstability across URL variants. To mitigate this, we introduce a component-aware canonicalization pipeline and a hybrid rule fusion framework, completely eliminating representation-based class flips and reducing the false positive rate (FPR) on secondarybenchmarks to 0.6%. However, to evaluate true operational readiness, we enforced a strict, hash-verified, version-controlled precollection freeze protocol utilizing domain-disjoint post-freeze external snapshots (OpenPhish and a post-freeze Cisco UmbrellaTop 1M snapshot). Despite our successful representation repair and strong secondary benchmarks, the frozen model exhibited amassive cross-source distribution shift, maintaining 91.34% phishing recall (95% CI: 87.4–94.8%) but degrading to an 86.84%FPR (95% CI: 83.6–90.0%) on independently sourced Umbrella benign observations. The combined ROC-AUC was 0.460 (95%CI: 0.392–0.524), providing no evidence of useful ranking discrimination under this source composition. Post hoc error analysisrevealed that, while complex subdomains and infrastructure reverse-DNS records were heavily enriched failure modes, the primarydriver was a broad vocabulary and score distribution shift across the unseen benign source. Our findings rigorously demonstrate thateliminating representation shortcuts is necessary, but insufficient, for cross-source generalization, highlighting a critical limitationin relying on constrained offline accuracy for operational deployment assessments.

Zenodo (CERN European Organization for Nuclear Research)
Tribhuvan University (NP)
Spam and Phishing Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.