Handcrafted features and ensemble learning for religious aggression detection in Bangla social media
Religious aggression on social media poses a growing threat to digital harmony and societal stability. In Bangladesh, where religion is deeply interwoven with cultural identity, online platforms often amplify intolerant and hostile discourse, fueling offline conflict. Automated detection of such aggression is therefore both a computational necessity and a social imperative. While hate speech detection in English has advanced considerably, Bangla, spoken by over 230 million people, remains critically under-resourced due to the lack of annotated corpora, linguistic complexity, and dialectal diversity. To address this gap, we introduce the Bangla Religious Aggression Comments (BRAC) dataset, comprising 20,000 manually annotated social media comments. Each comment is labeled along two dimensions: (i) whether it expresses religious aggression, and (ii) the specific religion targeted. Accordingly, we perform two sequential tasks: first, a binary classification to detect whether a comment is aggressive or non-aggressive, and second, a multi-class classification to identify the target religion, categorized as Muslim-AG (MAG), Hindu-AG (HAG), Christian-AG (CAG), and Buddhist-AG (BAG). We present the first comprehensive benchmark for this task, evaluating traditional machine learning, deep learning, and Transformer-based models. The proposed model is a soft-voting ensemble of XGBoost and Random Forest (XGB + RF) trained on TF–IDF features augmented with religion-word and offensive-word lexicon features. For aggression detection, the proposed ensemble achieves the highest accuracy (96.73%), while BanglaBERT attains the highest precision (95.55%) and F1-score (96.06%). For target religion classification, performance is comparable: the ensemble is slightly ahead in accuracy (94.67%), whereas BanglaBERT obtains the best F1-score (94.81%). These findings demonstrate complementary strengths between sparse lexical ensembles and contextual Transformer representations for domain-specific aggression detection in morphologically rich, low-resource languages such as Bangla.
Authors
- Arat Ibne Golam Mowla
- Riad Hossain (ORCID: https://orcid.org/0009-0005-6722-7140)
- Ayesha Banu
- Abid Hossain
- Mohammad Morshed Rana
Institutions
- Chittagong University of Engineering & Technology (BD)
- East Delta University (BD)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-10-01
- DOI
- https://doi.org/10.1007/s44163-026-02215-x
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00