Ethical and Socially Aware English–Norwegian Hate Speech Detection: A Systematic Design Framework
The rapid growth of social media has intensified the spread of hate speech targeting individuals based on race, gender, religion, and sexual orientation, creating significant technical, social, and ethical challenges. To address this, we propose a comprehensive English–Norwegian framework that unifies multi-class, multi-label, and multi-task learning for hate speech detection while embedding fairness, transparency, and accountability. The framework comprises six phases, including ethical system design, English–Norwegian data collection, bias analysis and mitigation, explainable AI, and stakeholder validation. A 14K-instance English–Norwegian dataset was constructed from public social media posts and enriched with identity-specific attributes (religion, race, gender, and sexual orientation) using human-in-the-loop and model-assisted annotation. The modeling approach applies multi-label binary fine-tuning to detect potentially hateful, offensive, aggressive, and neutral content, alongside multi-task learning for mutually exclusive demographic attributes. To promote ethical AI, identity-related prediction sensitivity is addressed through synonym-based augmentation, identity-conditioned label-contrast augmentation, and balanced loss re-weighting. Transparency is ensured using LIME to provide token-level explanations of model decisions. On the combined held-out evaluation partition, NorBERT-small achieved a macro F1-score of 0.92 and multi-label subset accuracy of 0.94 for toxicity classification. Language-stratified analysis further yielded macro F1-scores of 0.92 for English-only, 0.89 for Norwegian-only, and 0.87 for English–Norwegian code-switched text, providing a more detailed assessment of performance across the different language conditions, and high aggregate F1-scores for social-attribute extraction; however, these aggregate results should be interpreted cautiously because several social-attribute classes are sparsely represented. Compared with the baseline model, the mitigation stage reduces prediction sensitivity to many of the evaluated identity substitutions, although residual disparities remain for some identity pairs. The Machine Learning (ML) and Deep Learning (DL) models are included as conventional reference baselines. Because the evaluated model families employ different architectures and task-specific training procedures, the cross-family results are interpreted descriptively rather than as evidence of the intrinsic superiority of any particular model family. The reported findings are specific to the evaluated corpus; cross-platform, cross-domain, and external-dataset generalization were not evaluated and are therefore not claimed.
Authors
- Sule Yildirim Yayilgan (ORCID: https://orcid.org/0000-0002-1982-6609)
- Ehtesham Hashmi (ORCID: https://orcid.org/0009-0000-2526-9899)
Institutions
- Norwegian University of Science and Technology (NO)
Publication Details
- Journal
- Machine Learning and Knowledge Extraction
- Published
- 2026-09-21
- DOI
- https://doi.org/10.3390/make8090292
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00