Can Adversarial Attacks Target Cutting‐Edge Language Models Fine‐Tuned for Hate Speech and Toxicity Detection?—A State‐of‐the‐Art Evaluation and Analysis

ABSTRACT The advent of the transformer architecture and its different variants have led to tremendous improvement in several fields lying under the domain of NLP. However, recent studies have conclusively demonstrated that these models are vulnerable to adversarial attacks. At the same time, studies related to attacks on language models fine‐tuned for tasks with real‐world security implications are still limited. One such task is automated content moderation, that is, the detection and removal of hateful, offensive, or toxic remarks online. Though the automation of this task provides several advantages, primarily the ability to handle humongous volumes of data, it comes with the disadvantage of being vulnerable to adversarial manipulations. This study serves to perform adversarial experiments on a diverse set of transformer models fine‐tuned on two offensive and toxic content detection datasets. By catering to a diverse set of models differing in architectural design, pre‐training objectives and computational efficacy, this paper seeks to provide a comprehensive evaluation and analysis of adversarial robustness in the context of hate speech and toxicity detection. The presented findings reveal deep‐rooted concerns regarding the reliability and security of these models in real‐world applications, thereby underscoring the critical need to enhance their robustness.

Authors

Institutions

Publication Details

Journal
Concurrency and Computation Practice and Experience
Published
2026-09-18
DOI
https://doi.org/10.1002/cpe.70958
Primary Topic
Hate Speech and Cyberbullying Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Can Adversarial Attacks Target Cutting‐Edge Language Models Fine‐Tuned for Hate Speech and Toxicity Detection?—A State‐of‐the‐Art Evaluation and Analysis

Sajal Aggarwal, Dinesh Kumar Vishwakarma
Concurrency and Computation Practice and Experience
Hate Speech and Cyberbullying Detection
article

Can Adversarial Attacks Target Cutting‐Edge Language Models Fine‐Tuned for Hate Speech and Toxicity Detection?—A State‐of‐the‐Art Evaluation and Analysis

Sajal Aggarwal, Dinesh Kumar Vishwakarma
article en

Abstract

ABSTRACT The advent of the transformer architecture and its different variants have led to tremendous improvement in several fields lying under the domain of NLP. However, recent studies have conclusively demonstrated that these models are vulnerable to adversarial attacks. At the same time, studies related to attacks on language models fine‐tuned for tasks with real‐world security implications are still limited. One such task is automated content moderation, that is, the detection and removal of hateful, offensive, or toxic remarks online. Though the automation of this task provides several advantages, primarily the ability to handle humongous volumes of data, it comes with the disadvantage of being vulnerable to adversarial manipulations. This study serves to perform adversarial experiments on a diverse set of transformer models fine‐tuned on two offensive and toxic content detection datasets. By catering to a diverse set of models differing in architectural design, pre‐training objectives and computational efficacy, this paper seeks to provide a comprehensive evaluation and analysis of adversarial robustness in the context of hate speech and toxicity detection. The presented findings reveal deep‐rooted concerns regarding the reliability and security of these models in real‐world applications, thereby underscoring the critical need to enhance their robustness.

Concurrency and Computation Practice and ExperienceVol. 38(19)
Delhi Technological University (IN)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Hate Speech and Cyberbullying Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.