Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts
In this study, harmful content detection in Turkish texts is addressed as a three-class text classification problem comprising Harmless, Profanity and Sexist Discourse. For this purpose, open-access Turkish datasets and data compiled within the scope of the TÜBİTAK 1001 project were combined, and after removing cross-source duplicates and conflicting duplicates, a final dataset of 12051 texts was used. In the experimental process, traditional machine learning models and Transformer-based language models were compared under the same dataset, the same class structure and a 5-fold group-based stratified cross-validation protocol. For the traditional models, in addition to word- and character-level TF-IDF representations, Word2Vec- and FastText-based representations were used to evaluate Linear SVM, Logistic Regression, Decision Tree, Random Forest, Extra Trees and XGBoost algorithms. For the Transformer-based models, BERTurk, Electra-TR, ConvBERTurk and XLM-RoBERTa were adapted to the three-class classification task. Among the traditional models, the highest performance was obtained with Logistic Regression on TF-IDF representations (82.82% macro F1-score); Linear SVM produced very close results, while FastText- and Word2Vec-based representations remained comparatively lower. The obtained results showed that Transformer-based models offer higher and more balanced performance than traditional machine learning models. Among all models, the highest success was achieved with the ConvBERTurk model, which reached a 90.93% macro F1-score, 90.93% accuracy and a 98.09% ROC-AUC value. The BERTurk model exhibited a performance very close to ConvBERTurk with a 90.28% macro F1-score. A paired t-test over the five folds shows that this difference is not statistically significant (t(4) = 2.36, P=0.078). These findings indicate that contextual language representations are effective in Turkish harmful content classification. The study contributes to the literature by addressing Turkish harmful content detection as a multi-class problem that also includes the distinction between profanity and sexist discourse, and by comparing traditional models and Transformer-based models under the same leakage-controlled experimental protocol.
Authors
- Caner Balım (ORCID: https://orcid.org/0000-0002-1010-129X)
- Naim Karasekreter (ORCID: https://orcid.org/0000-0003-2892-6430)
- Özkan Aslan (ORCID: https://orcid.org/0000-0002-2680-5419)
Institutions
- Afyon Kocatepe University (TR)
Publication Details
- Journal
- Black Sea Journal of Engineering and Science
- Published
- 2026-09-14
- DOI
- https://doi.org/10.34248/bsengineering.1996841
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00