A large-scale explainable framework for Bangla mental health classification using synthetic data generation, comprehensive data analysis, and deep learning–assisted LinearSVC
In low-resource settings like Bangladesh where remains huge treatment gaps on mental health disorders because of stigma, poor infrastructure, and a shortage of trained professionals that indicates a critical health concern. The majority of current research is focused on high resource languages, leaving Bangla severely understudied due to severe data scarcity and privacy restrictions. However, recent advances in natural language processing (NLP) have made it possible to automatically analyze mental health-related textual expressions from textual data. To tackle these challenges, this work proposes a scalable, understandable, and ethically sound system for categorizing Bangla mental health texts. A large-scale, entirely synthetic, privacy-preserving dataset with severity annotations for about 15 million Bangla text instances spanning 15 mental disorder categories is presented in this work. The dataset is designed to overcome data constraints and guarantee reproducibility and ethical compliance. Significant performance improvements over previous Bangla mental health classification research are shown by experimental data. In particular, the suggested model gets a Macro F1-score of 0.9995 and an accuracy of 99.96% demonstrating strong performance compared with the evaluated baseline approaches. An additional mixed-data evaluation containing both real and synthetic samples was conducted to further assess robustness that achieved approximately 93% accuracy. Across all classes, precision and recall are consistently close to perfect, demonstrating strong and well-rounded performance. In order to assess the methodology’s usefulness, under more realistic linguistic conditions, additional experiments involving mixed real and synthetic Bangla text samples further indicate the practical applicability of the proposed framework. Extensive exploratory data analysis, calibration-aware evaluation, and class-wise error diagnostics further highlight the resilience and reliability of the proposed system. SHAP and LIME are used to further test model predictions in order to provide both global and instance-level interpretability. This work presents a comprehensive framework thereby providing a reproducible benchmark for low-resource mental health research that integrates explainable modeling, validated synthetic data generation, and additional mixed-data evaluation while maintaining transparency and computational efficiency for large-scale Bangla mental health NLP.
Authors
- Mahfuzulhoq Chowdhury (ORCID: https://orcid.org/0000-0002-3006-4596)
- Payel Sen
Institutions
- Chittagong University of Engineering & Technology (BD)
Publication Details
- Journal
- Discover Applied Sciences
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1007/s42452-026-09362-x
- Primary Topic
- Mental Health via Writing
- Type
- article
- Field-Weighted Citation Impact
- 0.00