Evasion-hate: targeted augmentation for hate speech detection under prompt-induced surface-form shifts
Purpose Online hate speech detection is increasingly used as a decision-support component in content moderation, but most detectors are evaluated on clean held-out text. In practice, users can preserve harmful intent while changing surface form through spelling obfuscation, coded wording or code-switching, creating a robustness gap between benchmark performance and deployment reliability. This paper introduces Evasion-Hate, a targeted, quality-filtered augmentation pipeline for three-way hate/offensive/normal classification in English and Vietnamese. Design/methodology/approach The pipeline generates synthetic large language model (LLM) candidate rewrites, filters them with checks independent of the downstream classifiers under evaluation and evaluates robustness with paired bootstrap testing, matched quantity controls, competitive augmentation baselines and external functional suites. Findings Under the main evasion protocol, Evasion-Hate achieves the strongest English robustness result among comparable three-way systems evaluated under our protocol: XLM-R robust Macro-F1 rises from 0.4783 to 0.7582 (+27.99 pp); model-specific PFR is 37.16% for the baseline and 16.66% for Evasion-Hate, while model-specific harmful-only ASR is 56.56% and 11.57%, respectively. Vietnamese models show statistically supported gains on the reconstructed synthetic evaluation sets (PhoBERT-v2 +4.60 pp; ViSoBERT +3.29 pp; both significant by pooled paired bootstrap), although the Vietnamese transform families remain below or near the pre-specified fidelity threshold and should be interpreted diagnostically. Matched quantity controls show the gain is not explained by additional training volume; external HateCheck spelling-variation flagged rate improves by 26.42 pp. Leave-one-transformation-out evaluation shows partial and family-dependent transfer. In English, holding out euphemistic and coded rewriting yields a +18.01 pp gain that remains supported after Holm correction, whereas the +8.48 pp gain for held-out code-switching does not survive multiplicity correction and the character-level family shows little change (+0.31 pp). In Vietnamese, held-out code-switching shows a positive +2.36 pp effect with a bootstrap interval excluding zero, but it does not survive Holm correction; the remaining folds and the pooled evaluation show no supported cross-family transfer. Originality/value The results establish Evasion-Hate as an auditable protocol for targeted evasion robustness. Cross-family evaluation further identifies where robustness transfers to unseen transformations and where explicit family coverage remains necessary.
Authors
- Thien Khai Tran (ORCID: https://orcid.org/0000-0003-2119-7167)
- Lê Mạnh Thạnh (ORCID: https://orcid.org/0009-0006-3936-4858)
Institutions
- University of Economics Ho Chi Minh City (VN)
- University of Finance - Marketing (VN)
Publication Details
- Journal
- Data Technologies and Applications
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1108/dta-08-2026-1078
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00