Hybrid Transformer Architecture for Context-Aware Spam Email Classification
Spam and phishing email remain a persistent cybersecurity threat and a primary delivery vector for credential theft, malware, and ransomware. Detection accuracies above 99% are routinely reported. Such figures are typically obtained on small, single-source, monolingual corpora in which near-duplicate campaign templates span the training and test partitions. This study re-examines what they measure. A corpus of 99,707 emails is assembled from four independent public collections spanning multiple languages, campaign-aware deduplication is applied before splitting, and nine models—classical, convolutional, Transformer, and multimodal—are evaluated under two protocols: an in-distribution split and a cross-source split in which the evaluation source is withheld entirely from training. The multimodal model fine-tunes DistilBERT and fuses its contextual embeddings with thirteen structural features into a 781-dimensional representation classified by a multilayer perceptron. In distribution, six of the nine models fall within one F1 point of each other (ΔF1 = 0.0001, 95% CI [−0.0012, 0.0012], fusion vs. the fine-tuned text-only baseline). Against an identical classification head without the structural block—the controlled comparison isolating the structural contribution—fusion produces a small but statistically significant gain in distribution (ΔF1 = 0.0014, 95% CI [0.0005, 0.0023]; Δrecall = 0.0030, CI [0.0016, 0.0044]) that grows roughly five-and-a-half-fold under cross-source evaluation (ΔF1 = 0.0078, CI [0.0036, 0.0120]; Δrecall = 0.0107, CI [0.0031, 0.0180]). The contribution of multimodal fusion therefore scales with distribution shift rather than appearing only under it and is largest exactly where deployment conditions differ most from training. Against the strongest classical baseline (a linear SVM), the fusion model attains a false-positive rate less than one-fifth as large under cross-source evaluation (0.91% vs. 5.40%) alongside significantly higher F1 (0.9488 vs. 0.9188, non-overlapping bootstrap CIs); in distribution, it attains marginally higher AUC (0.9990 vs. 0.9985) at a marginally higher false-positive rate (0.94% vs. 0.37%). SHAP places 97.6% of signal in the text embeddings, yet that residual share is what produces the cross-source gain—attribution magnitude measured in distribution does not predict conditional contribution under shift.
Authors
- Hamid Jahankhani (ORCID: https://orcid.org/0000-0002-8288-4609)
- Umair B. Chaudhry (ORCID: https://orcid.org/0000-0002-8609-8357)
- Geldi Xhafaj
Institutions
- Queen Mary University of London (GB)
- Northumbria University (GB)
Publication Details
- Journal
- Future Internet
- Published
- 2026-10-09
- DOI
- https://doi.org/10.3390/fi18100542
- Primary Topic
- Spam and Phishing Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00