SUICIDE AND CRISIS RISK DETECTION IN MENTAL HEALTH CHATBOTS: TECHNICAL APPROACHES, SAFETY EVIDENCE, AND GOVERNANCE CHALLENGES

Mental health chatbots built on large language models are now used at population scale, including by adolescents disclosing suicidal thoughts, yet the evidence base supporting them was generated to measure symptom change rather than crisis safety. This narrative review synthesises 43 records published between 2021 and 2026 and identified through PubMed, Scopus and Web of Science, organised around technical approaches to suicide and crisis risk detection, empirical safety evidence, and governance responses. Detection from natural language is technically feasible: classifiers applied to crisis conversation corpora reach areas under the curve between 0.89 and 0.92, and prompt-defined thresholds can reduce false negatives to zero at sub-second latency, although only at the cost of substantially elevated false alarms. Reported performance is not comparable across studies because label provenance varies, and model errors concentrate on the cases about which expert clinicians themselves disagree. Deployed systems perform poorly: among 29 commercial applications tested against escalating suicidal risk scenarios, none met the criteria for an adequate response, and general-purpose models align with clinical judgement at the extremes of the risk continuum but not in the intermediate range where assessment is most consequential. Where crisis handling has demonstrably worked, detection was separated from the conversational model and coupled to a human escalation pathway. Effectiveness reviews rarely measure safety, and where safety was examined, adverse events including new-onset suicidal tendency were recorded. Governance should mandate crisis detection and referral, close classification gaps between medical device and general-purpose AI regulation, and require continuous independent auditing.

Authors

Institutions

Publication Details

Journal
International Journal of Innovative Technologies in Social Science
Published
2026-09-16
DOI
https://doi.org/10.31435/ijitss.3(51).2026.6484
Primary Topic
Digital Mental Health Interventions
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SUICIDE AND CRISIS RISK DETECTION IN MENTAL HEALTH CHATBOTS: TECHNICAL APPROACHES, SAFETY EVIDENCE, AND GOVERNANCE CHALLENGES

Karol Wrzosek, Maria Leśniak, P. Kowalewski, Karina Udrycka et al.
International Journal of Innovative Technologies in Social Science
Digital Mental Health Interventions
article

SUICIDE AND CRISIS RISK DETECTION IN MENTAL HEALTH CHATBOTS: TECHNICAL APPROACHES, SAFETY EVIDENCE, AND GOVERNANCE CHALLENGES

Karol Wrzosek, Maria Leśniak, P. Kowalewski, Karina Udrycka, Kornelia Kuchta, Jan Poleszak
article en

Abstract

Mental health chatbots built on large language models are now used at population scale, including by adolescents disclosing suicidal thoughts, yet the evidence base supporting them was generated to measure symptom change rather than crisis safety. This narrative review synthesises 43 records published between 2021 and 2026 and identified through PubMed, Scopus and Web of Science, organised around technical approaches to suicide and crisis risk detection, empirical safety evidence, and governance responses. Detection from natural language is technically feasible: classifiers applied to crisis conversation corpora reach areas under the curve between 0.89 and 0.92, and prompt-defined thresholds can reduce false negatives to zero at sub-second latency, although only at the cost of substantially elevated false alarms. Reported performance is not comparable across studies because label provenance varies, and model errors concentrate on the cases about which expert clinicians themselves disagree. Deployed systems perform poorly: among 29 commercial applications tested against escalating suicidal risk scenarios, none met the criteria for an adequate response, and general-purpose models align with clinical judgement at the extremes of the risk continuum but not in the intermediate range where assessment is most consequential. Where crisis handling has demonstrably worked, detection was separated from the conversational model and coupled to a human escalation pathway. Effectiveness reviews rarely measure safety, and where safety was examined, adverse events including new-onset suicidal tendency were recorded. Governance should mandate crisis detection and referral, close classification gaps between medical device and general-purpose AI regulation, and require continuous independent auditing.

International Journal of Innovative Technologies in Social ScienceVol. 3(3(51))
Medical University of Lublin (PL), Medical University of Lodz (PL)
Openalex Percentile: Top 9%
Digital Mental Health Interventions
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.