Current Developments and Future Trends in Turkish Hate Speech Detection: A Systematic Review

Although Turkish is a linguistically distinct, low-resource language, its popularity on social media makes it important to the HSD literature. This review demonstrates which TML, DL and emerging Transformer-based models are used in THSD research and for which reason. It systematically analyses the datasets used in these studies. Furthermore, identifies emerging directions, including the integration of XAI frameworks, multi-modal (text, image, audio, video, etc.) constraints and synthetic data creation. A PRISMA-guided search was conducted in Scopus, IEEE Xplore, Wiley, Science Direct and Springer databases for publications dated from 1 January 2022 to 31 August 2026. Of the 56 reports assessed in full text, 41 studies were included. Dataset characteristics, preprocessing methods, labeling schemes and evaluation protocols were reported narratively. Class imbalance was identified in commonly used datasets. Homophobic comments account for 1226 of the 31,290 HATC examples (3.9%); the OffensEval-2020 Turkish dataset contains only 6848 offensive data points out of 35,284 samples. Different data allocations were detected for the same datasets. The Augmented Offensive Language Dataset (53,005 samples) corpus reported here has different train/validation/test allocations: 31,803/10,601/10,601 and 42,398/1756/8851. Differences in labels, preprocessing, corpus versions and evaluation protocols preclude a reliable overall ranking of the reported models. This review presents current approaches and their limitations. It also gives a roadmap for next-generation solutions related to the Turkish language.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-10-06
DOI
https://doi.org/10.3390/app16199899
Primary Topic
Hate Speech and Cyberbullying Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Current Developments and Future Trends in Turkish Hate Speech Detection: A Systematic Review

Muhammed Ali Aydın, Mustafa Kara, Hasan Hüseyin Balık, Berkay Özçam et al.
Applied Sciences
Hate Speech and Cyberbullying Detection
article

Current Developments and Future Trends in Turkish Hate Speech Detection: A Systematic Review

Muhammed Ali Aydın, Mustafa Kara, Hasan Hüseyin Balık, Berkay Özçam, Atakan Akgül, Üsame Yiğit
article en

Abstract

Although Turkish is a linguistically distinct, low-resource language, its popularity on social media makes it important to the HSD literature. This review demonstrates which TML, DL and emerging Transformer-based models are used in THSD research and for which reason. It systematically analyses the datasets used in these studies. Furthermore, identifies emerging directions, including the integration of XAI frameworks, multi-modal (text, image, audio, video, etc.) constraints and synthetic data creation. A PRISMA-guided search was conducted in Scopus, IEEE Xplore, Wiley, Science Direct and Springer databases for publications dated from 1 January 2022 to 31 August 2026. Of the 56 reports assessed in full text, 41 studies were included. Dataset characteristics, preprocessing methods, labeling schemes and evaluation protocols were reported narratively. Class imbalance was identified in commonly used datasets. Homophobic comments account for 1226 of the 31,290 HATC examples (3.9%); the OffensEval-2020 Turkish dataset contains only 6848 offensive data points out of 35,284 samples. Different data allocations were detected for the same datasets. The Augmented Offensive Language Dataset (53,005 samples) corpus reported here has different train/validation/test allocations: 31,803/10,601/10,601 and 42,398/1756/8851. Differences in labels, preprocessing, corpus versions and evaluation protocols preclude a reliable overall ranking of the reported models. This review presents current approaches and their limitations. It also gives a roadmap for next-generation solutions related to the Turkish language.

Applied SciencesVol. 16(19)
Yıldız Technical University (TR), Milli Savunma Üniversitesi (TR), Istanbul University-Cerrahpaşa (TR), Istanbul Technical University (TR), Istanbul University (TR), Turkish Air Force Academy (TR)
Openalex Percentile: Top 11%
Hate Speech and Cyberbullying Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.