A large-scale explainable framework for Bangla mental health classification using synthetic data generation, comprehensive data analysis, and deep learning–assisted LinearSVC

In low-resource settings like Bangladesh where remains huge treatment gaps on mental health disorders because of stigma, poor infrastructure, and a shortage of trained professionals that indicates a critical health concern. The majority of current research is focused on high resource languages, leaving Bangla severely understudied due to severe data scarcity and privacy restrictions. However, recent advances in natural language processing (NLP) have made it possible to automatically analyze mental health-related textual expressions from textual data. To tackle these challenges, this work proposes a scalable, understandable, and ethically sound system for categorizing Bangla mental health texts. A large-scale, entirely synthetic, privacy-preserving dataset with severity annotations for about 15 million Bangla text instances spanning 15 mental disorder categories is presented in this work. The dataset is designed to overcome data constraints and guarantee reproducibility and ethical compliance. Significant performance improvements over previous Bangla mental health classification research are shown by experimental data. In particular, the suggested model gets a Macro F1-score of 0.9995 and an accuracy of 99.96% demonstrating strong performance compared with the evaluated baseline approaches. An additional mixed-data evaluation containing both real and synthetic samples was conducted to further assess robustness that achieved approximately 93% accuracy. Across all classes, precision and recall are consistently close to perfect, demonstrating strong and well-rounded performance. In order to assess the methodology’s usefulness, under more realistic linguistic conditions, additional experiments involving mixed real and synthetic Bangla text samples further indicate the practical applicability of the proposed framework. Extensive exploratory data analysis, calibration-aware evaluation, and class-wise error diagnostics further highlight the resilience and reliability of the proposed system. SHAP and LIME are used to further test model predictions in order to provide both global and instance-level interpretability. This work presents a comprehensive framework thereby providing a reproducible benchmark for low-resource mental health research that integrates explainable modeling, validated synthetic data generation, and additional mixed-data evaluation while maintaining transparency and computational efficiency for large-scale Bangla mental health NLP.

Authors

Institutions

Publication Details

Journal
Discover Applied Sciences
Published
2026-09-21
DOI
https://doi.org/10.1007/s42452-026-09362-x
Primary Topic
Mental Health via Writing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A large-scale explainable framework for Bangla mental health classification using synthetic data generation, comprehensive data analysis, and deep learning–assisted LinearSVC

Mahfuzulhoq Chowdhury, Payel Sen
Discover Applied Sciences
Mental Health via Writing
article

A large-scale explainable framework for Bangla mental health classification using synthetic data generation, comprehensive data analysis, and deep learning–assisted LinearSVC

Mahfuzulhoq Chowdhury, Payel Sen
article en

Abstract

In low-resource settings like Bangladesh where remains huge treatment gaps on mental health disorders because of stigma, poor infrastructure, and a shortage of trained professionals that indicates a critical health concern. The majority of current research is focused on high resource languages, leaving Bangla severely understudied due to severe data scarcity and privacy restrictions. However, recent advances in natural language processing (NLP) have made it possible to automatically analyze mental health-related textual expressions from textual data. To tackle these challenges, this work proposes a scalable, understandable, and ethically sound system for categorizing Bangla mental health texts. A large-scale, entirely synthetic, privacy-preserving dataset with severity annotations for about 15 million Bangla text instances spanning 15 mental disorder categories is presented in this work. The dataset is designed to overcome data constraints and guarantee reproducibility and ethical compliance. Significant performance improvements over previous Bangla mental health classification research are shown by experimental data. In particular, the suggested model gets a Macro F1-score of 0.9995 and an accuracy of 99.96% demonstrating strong performance compared with the evaluated baseline approaches. An additional mixed-data evaluation containing both real and synthetic samples was conducted to further assess robustness that achieved approximately 93% accuracy. Across all classes, precision and recall are consistently close to perfect, demonstrating strong and well-rounded performance. In order to assess the methodology’s usefulness, under more realistic linguistic conditions, additional experiments involving mixed real and synthetic Bangla text samples further indicate the practical applicability of the proposed framework. Extensive exploratory data analysis, calibration-aware evaluation, and class-wise error diagnostics further highlight the resilience and reliability of the proposed system. SHAP and LIME are used to further test model predictions in order to provide both global and instance-level interpretability. This work presents a comprehensive framework thereby providing a reproducible benchmark for low-resource mental health research that integrates explainable modeling, validated synthetic data generation, and additional mixed-data evaluation while maintaining transparency and computational efficiency for large-scale Bangla mental health NLP.

Discover Applied Sciences
Chittagong University of Engineering & Technology (BD)
Openalex Percentile: Top 7%
Mental Health via Writing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.