Can Adversarial Attacks Target Cutting‐Edge Language Models Fine‐Tuned for Hate Speech and Toxicity Detection?—A State‐of‐the‐Art Evaluation and Analysis
ABSTRACT The advent of the transformer architecture and its different variants have led to tremendous improvement in several fields lying under the domain of NLP. However, recent studies have conclusively demonstrated that these models are vulnerable to adversarial attacks. At the same time, studies related to attacks on language models fine‐tuned for tasks with real‐world security implications are still limited. One such task is automated content moderation, that is, the detection and removal of hateful, offensive, or toxic remarks online. Though the automation of this task provides several advantages, primarily the ability to handle humongous volumes of data, it comes with the disadvantage of being vulnerable to adversarial manipulations. This study serves to perform adversarial experiments on a diverse set of transformer models fine‐tuned on two offensive and toxic content detection datasets. By catering to a diverse set of models differing in architectural design, pre‐training objectives and computational efficacy, this paper seeks to provide a comprehensive evaluation and analysis of adversarial robustness in the context of hate speech and toxicity detection. The presented findings reveal deep‐rooted concerns regarding the reliability and security of these models in real‐world applications, thereby underscoring the critical need to enhance their robustness.
Authors
- Sajal Aggarwal (ORCID: https://orcid.org/0000-0002-8261-2662)
- Dinesh Kumar Vishwakarma (ORCID: https://orcid.org/0000-0002-1026-0047)
Institutions
- Delhi Technological University (IN)
Publication Details
- Journal
- Concurrency and Computation Practice and Experience
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1002/cpe.70958
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00