Fine-tuning large language models for binary and multiclass target-specific hate speech detection
Abstract Social media platforms are widely used by individuals from diverse demographic backgrounds to share their opinions and daily experiences. The anonymity provided by these platforms has increased the prevalence of hate speech, causing significant psychological harm to targeted groups. Existing hate speech detection approaches predominantly rely on binary classification, which fails to capture the lexical and contextual diversity of hate categories, leading to the misclassification of hate against minority groups. This study addresses this limitation by evaluating the effectiveness of five large language models (LLMs) for target-specific multiclass hate speech detection across six classes. The evaluation uses a dataset of 12,957 tweets, including 5,455 labeled hate samples across five categories and 7,502 non-hate samples. The experiments compare zero-shot, few-shot, and fine-tuning methods and evaluate LLMs against replicated baseline approaches. Fine-tuning outperformed zero-shot and few-shot settings with the fine-tuned DeepSeek model, achieving the best multiclass performance with macro and weighted F1 scores of 59.91 and 76.71%, respectively. LLM fine-tuning significantly improves multiclass hate speech detection performance and reduces the number of unclassified outputs. Compared with several LLMs, BERT, and replicated classical machine learning and deep learning baselines, the fine-tuned DeepSeek improved the weighted F1 score ranging from 2.59 to 29.5%. Evaluation on a remapped subset of the external ETHOS dataset yields an F1 score of 88.31%, demonstrating that the model retains robust classification behavior under an aligned label space. These results highlight the effectiveness of LLM-based adaptation for target-specific content moderation tasks, while minority-class representation remains a challenge for real-world deployment.
Authors
- Sanaa Kaddoura (ORCID: https://orcid.org/0000-0002-4384-4364)
- Sumaia A. Al‐Kohlani (ORCID: https://orcid.org/0000-0003-4586-0031)
Institutions
- United Arab Emirates University (AE)
- Zayed University (AE)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1007/s44163-026-02354-1
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00