The linguistic features of isiZulu cyberbullying: A pilot study
To combat the spread of fake news, hate speech and cyberbullying, social media platforms have increasingly turned to automatic detection tools, the effectiveness of which depends on the availability of robust linguistic datasets. This article reports on a pilot study that compiled a dataset of aggressive isiZulu language used in online communication. The study aimed to identify isiZulu words and phrases perceived as aggressive and to describe their linguistic and thematic features. Using an established taxonomy of cyberbullying indicators, the isiZulu dataset of n = 695 words was analysed with Anthony Concordance, a text analysis software package, to identify high-frequency words and keywords in context. The analysis revealed that the isiZulu dataset shares several features with existing cyberbullying corpora, including the frequent use of swear words, references to biological processes and body parts, animal terms and the second-person pronoun. While the limited dataset constrains the reliability, validity and generalisability of the findings, this pilot project represents an initial step toward developing a comprehensive isiZulu cyberbullying language dataset. As such, it contributes to broader initiatives in multilingual online safety and automated content moderation.
Authors
- Shamila Naidoo (ORCID: https://orcid.org/0000-0002-1013-2221)
- Thabiso Ntuli (ORCID: https://orcid.org/0000-0002-7504-5540)
Institutions
- University of KwaZulu-Natal (ZA)
Publication Details
- Journal
- Southern African Linguistics and Applied Language Studies
- Published
- 2026-10-05
- DOI
- https://doi.org/10.2989/16073614.2026.2694600
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00