Ensemble of Domain-Specific Natural Language Processing and Large Language Models for Detecting Suicidal Ideation and Mental Health Conditions in Social Media Text: Development and Evaluation Study
BACKGROUND Mental illness contributes substantially to global disability, and public adoption of AI for mental health support is accelerating without commensurate safety evaluation. General-purpose large language models can fail to recognize naturalistic text expressing suicidal ideation, with immediate clinical consequences. Domain-specific natural language processing methods offer a contrasting approach, but ensembles integrating the two have not been formally characterized. OBJECTIVE We aimed to (1) benchmark 5 modeling approaches for classifying mental health–related social media text, (2) develop and optimize ensembles integrating a domain-specific natural language processing classifier with a fine-tuned large language model, and (3) derive a closed-form framework constraining ensemble weighting in safety-critical classification. METHODS We analyzed 60,889 publicly available Reddit and Twitter texts spanning 9 categories (anxiety, bipolar disorder, depression, a normal baseline, personality disorder, stress, suicidal ideation, attention-deficit/hyperactivity disorder, and autism spectrum disorder). Labels were derived from the originating subreddit or a self-stated condition and are not clinical diagnoses. Five architectures were compared: prompt-engineered base GPT-4o-mini; fine-tuned GPT-4o-mini; a domain-specific classifier (support vector machine over symptom-informed lexical features); and 2 ensembles of these models, hybrid probability–indicator fusion and soft probability fusion. Data were split into 70%, 10%, and 20% at the post level (n=42,622, n=6089, and n=12,178); weights were selected on the validation split and all reported performance comes from the held-out test split. Closed-form bounds on permissible language model weights were derived for both strategies. CIs are paired bootstrap intervals (n=10,000 replicates); accuracy differences were tested with exact McNemar tests. RESULTS The domain-specific classifier achieved 91.9% accuracy (95% CI 91.4%-92.4%), exceeding the base (58.7%) and fine-tuned (88.6%) large language models. Both ensembles outperformed either constituent model, reaching 93.9% (soft fusion; 95% CI 93.5%-94.3%) and 93.8% (hybrid fusion; 95% CI 93.4%-94.2%) at a weighting of 60% classifier to 40% fine-tuned model, selected on validation (P<.001 for both against the classifier alone; the 2 ensembles did not differ from each other [P=.16]). The hybrid ensemble reduced the miss rate for the suicidal ideation category from 32.5% (650/2000) to 1.9% (37/2000), a 17.6-fold reduction relative to the base model, although its advantage over the classifier alone within that category was not significant (P=.68). Accuracy fell discontinuously above a 50% language model weight, where the hybrid ensemble reduced exactly to the fine-tuned model; the selected weight satisfies the derived instance-level bound of 44.0%. CONCLUSIONS Domain-specific clinical grounding remained necessary for stable ensemble performance in this corpus, and probabilistic fusion outperformed both constituent models when language model weight was bounded by a derivable, model-agnostic constraint. Because labels were community-inferred and only one corpus was analyzed, these results are hypotheses requiring external validation on clinically characterized data before use in decision support.
Authors
- Andrea Johansson Capusan (ORCID: https://orcid.org/0000-0003-1758-2206)
- Thomas Kallstenius
- Adam Williamson (ORCID: https://orcid.org/0000-0002-5632-0084)
- Basil Duvernoy (ORCID: https://orcid.org/0000-0001-5429-5594)
- Adam Kallstenius (ORCID: https://orcid.org/0009-0000-9662-7550)
Publication Details
- Journal
- JMIR Mental Health
- Published
- 2026-10-08
- DOI
- https://doi.org/10.2196/98925
- Primary Topic
- Mental Health via Writing
- Type
- article
- Field-Weighted Citation Impact
- 0.00