Ethical and Socially Aware English–Norwegian Hate Speech Detection: A Systematic Design Framework

The rapid growth of social media has intensified the spread of hate speech targeting individuals based on race, gender, religion, and sexual orientation, creating significant technical, social, and ethical challenges. To address this, we propose a comprehensive English–Norwegian framework that unifies multi-class, multi-label, and multi-task learning for hate speech detection while embedding fairness, transparency, and accountability. The framework comprises six phases, including ethical system design, English–Norwegian data collection, bias analysis and mitigation, explainable AI, and stakeholder validation. A 14K-instance English–Norwegian dataset was constructed from public social media posts and enriched with identity-specific attributes (religion, race, gender, and sexual orientation) using human-in-the-loop and model-assisted annotation. The modeling approach applies multi-label binary fine-tuning to detect potentially hateful, offensive, aggressive, and neutral content, alongside multi-task learning for mutually exclusive demographic attributes. To promote ethical AI, identity-related prediction sensitivity is addressed through synonym-based augmentation, identity-conditioned label-contrast augmentation, and balanced loss re-weighting. Transparency is ensured using LIME to provide token-level explanations of model decisions. On the combined held-out evaluation partition, NorBERT-small achieved a macro F1-score of 0.92 and multi-label subset accuracy of 0.94 for toxicity classification. Language-stratified analysis further yielded macro F1-scores of 0.92 for English-only, 0.89 for Norwegian-only, and 0.87 for English–Norwegian code-switched text, providing a more detailed assessment of performance across the different language conditions, and high aggregate F1-scores for social-attribute extraction; however, these aggregate results should be interpreted cautiously because several social-attribute classes are sparsely represented. Compared with the baseline model, the mitigation stage reduces prediction sensitivity to many of the evaluated identity substitutions, although residual disparities remain for some identity pairs. The Machine Learning (ML) and Deep Learning (DL) models are included as conventional reference baselines. Because the evaluated model families employ different architectures and task-specific training procedures, the cross-family results are interpreted descriptively rather than as evidence of the intrinsic superiority of any particular model family. The reported findings are specific to the evaluated corpus; cross-platform, cross-domain, and external-dataset generalization were not evaluated and are therefore not claimed.

Authors

Institutions

Publication Details

Journal
Machine Learning and Knowledge Extraction
Published
2026-09-21
DOI
https://doi.org/10.3390/make8090292
Primary Topic
Hate Speech and Cyberbullying Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Ethical and Socially Aware English–Norwegian Hate Speech Detection: A Systematic Design Framework

Sule Yildirim Yayilgan, Ehtesham Hashmi
Machine Learning and Knowledge Extraction
Hate Speech and Cyberbullying Detection
article

Ethical and Socially Aware English–Norwegian Hate Speech Detection: A Systematic Design Framework

Sule Yildirim Yayilgan, Ehtesham Hashmi
article en

Abstract

The rapid growth of social media has intensified the spread of hate speech targeting individuals based on race, gender, religion, and sexual orientation, creating significant technical, social, and ethical challenges. To address this, we propose a comprehensive English–Norwegian framework that unifies multi-class, multi-label, and multi-task learning for hate speech detection while embedding fairness, transparency, and accountability. The framework comprises six phases, including ethical system design, English–Norwegian data collection, bias analysis and mitigation, explainable AI, and stakeholder validation. A 14K-instance English–Norwegian dataset was constructed from public social media posts and enriched with identity-specific attributes (religion, race, gender, and sexual orientation) using human-in-the-loop and model-assisted annotation. The modeling approach applies multi-label binary fine-tuning to detect potentially hateful, offensive, aggressive, and neutral content, alongside multi-task learning for mutually exclusive demographic attributes. To promote ethical AI, identity-related prediction sensitivity is addressed through synonym-based augmentation, identity-conditioned label-contrast augmentation, and balanced loss re-weighting. Transparency is ensured using LIME to provide token-level explanations of model decisions. On the combined held-out evaluation partition, NorBERT-small achieved a macro F1-score of 0.92 and multi-label subset accuracy of 0.94 for toxicity classification. Language-stratified analysis further yielded macro F1-scores of 0.92 for English-only, 0.89 for Norwegian-only, and 0.87 for English–Norwegian code-switched text, providing a more detailed assessment of performance across the different language conditions, and high aggregate F1-scores for social-attribute extraction; however, these aggregate results should be interpreted cautiously because several social-attribute classes are sparsely represented. Compared with the baseline model, the mitigation stage reduces prediction sensitivity to many of the evaluated identity substitutions, although residual disparities remain for some identity pairs. The Machine Learning (ML) and Deep Learning (DL) models are included as conventional reference baselines. Because the evaluated model families employ different architectures and task-specific training procedures, the cross-family results are interpreted descriptively rather than as evidence of the intrinsic superiority of any particular model family. The reported findings are specific to the evaluated corpus; cross-platform, cross-domain, and external-dataset generalization were not evaluated and are therefore not claimed.

Machine Learning and Knowledge ExtractionVol. 8(9)
Norwegian University of Science and Technology (NO)
Gender equality
Openalex Percentile: Top 8%
Hate Speech and Cyberbullying Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.