Hybrid Transformer Architecture for Context-Aware Spam Email Classification

Spam and phishing email remain a persistent cybersecurity threat and a primary delivery vector for credential theft, malware, and ransomware. Detection accuracies above 99% are routinely reported. Such figures are typically obtained on small, single-source, monolingual corpora in which near-duplicate campaign templates span the training and test partitions. This study re-examines what they measure. A corpus of 99,707 emails is assembled from four independent public collections spanning multiple languages, campaign-aware deduplication is applied before splitting, and nine models—classical, convolutional, Transformer, and multimodal—are evaluated under two protocols: an in-distribution split and a cross-source split in which the evaluation source is withheld entirely from training. The multimodal model fine-tunes DistilBERT and fuses its contextual embeddings with thirteen structural features into a 781-dimensional representation classified by a multilayer perceptron. In distribution, six of the nine models fall within one F1 point of each other (ΔF1 = 0.0001, 95% CI [−0.0012, 0.0012], fusion vs. the fine-tuned text-only baseline). Against an identical classification head without the structural block—the controlled comparison isolating the structural contribution—fusion produces a small but statistically significant gain in distribution (ΔF1 = 0.0014, 95% CI [0.0005, 0.0023]; Δrecall = 0.0030, CI [0.0016, 0.0044]) that grows roughly five-and-a-half-fold under cross-source evaluation (ΔF1 = 0.0078, CI [0.0036, 0.0120]; Δrecall = 0.0107, CI [0.0031, 0.0180]). The contribution of multimodal fusion therefore scales with distribution shift rather than appearing only under it and is largest exactly where deployment conditions differ most from training. Against the strongest classical baseline (a linear SVM), the fusion model attains a false-positive rate less than one-fifth as large under cross-source evaluation (0.91% vs. 5.40%) alongside significantly higher F1 (0.9488 vs. 0.9188, non-overlapping bootstrap CIs); in distribution, it attains marginally higher AUC (0.9990 vs. 0.9985) at a marginally higher false-positive rate (0.94% vs. 0.37%). SHAP places 97.6% of signal in the text embeddings, yet that residual share is what produces the cross-source gain—attribution magnitude measured in distribution does not predict conditional contribution under shift.

Authors

Institutions

Publication Details

Journal
Future Internet
Published
2026-10-09
DOI
https://doi.org/10.3390/fi18100542
Primary Topic
Spam and Phishing Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Hybrid Transformer Architecture for Context-Aware Spam Email Classification

Hamid Jahankhani, Umair B. Chaudhry, Geldi Xhafaj
Future Internet
Spam and Phishing Detection
article

Hybrid Transformer Architecture for Context-Aware Spam Email Classification

Hamid Jahankhani, Umair B. Chaudhry, Geldi Xhafaj
article en

Abstract

Spam and phishing email remain a persistent cybersecurity threat and a primary delivery vector for credential theft, malware, and ransomware. Detection accuracies above 99% are routinely reported. Such figures are typically obtained on small, single-source, monolingual corpora in which near-duplicate campaign templates span the training and test partitions. This study re-examines what they measure. A corpus of 99,707 emails is assembled from four independent public collections spanning multiple languages, campaign-aware deduplication is applied before splitting, and nine models—classical, convolutional, Transformer, and multimodal—are evaluated under two protocols: an in-distribution split and a cross-source split in which the evaluation source is withheld entirely from training. The multimodal model fine-tunes DistilBERT and fuses its contextual embeddings with thirteen structural features into a 781-dimensional representation classified by a multilayer perceptron. In distribution, six of the nine models fall within one F1 point of each other (ΔF1 = 0.0001, 95% CI [−0.0012, 0.0012], fusion vs. the fine-tuned text-only baseline). Against an identical classification head without the structural block—the controlled comparison isolating the structural contribution—fusion produces a small but statistically significant gain in distribution (ΔF1 = 0.0014, 95% CI [0.0005, 0.0023]; Δrecall = 0.0030, CI [0.0016, 0.0044]) that grows roughly five-and-a-half-fold under cross-source evaluation (ΔF1 = 0.0078, CI [0.0036, 0.0120]; Δrecall = 0.0107, CI [0.0031, 0.0180]). The contribution of multimodal fusion therefore scales with distribution shift rather than appearing only under it and is largest exactly where deployment conditions differ most from training. Against the strongest classical baseline (a linear SVM), the fusion model attains a false-positive rate less than one-fifth as large under cross-source evaluation (0.91% vs. 5.40%) alongside significantly higher F1 (0.9488 vs. 0.9188, non-overlapping bootstrap CIs); in distribution, it attains marginally higher AUC (0.9990 vs. 0.9985) at a marginally higher false-positive rate (0.94% vs. 0.37%). SHAP places 97.6% of signal in the text embeddings, yet that residual share is what produces the cross-source gain—attribution magnitude measured in distribution does not predict conditional contribution under shift.

Future InternetVol. 18(10)
Queen Mary University of London (GB), Northumbria University (GB)
Openalex Percentile: Top 6%
Spam and Phishing Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.