The Responder and the Supervisor: Trajectory-Level Safety Control for Mental-Health LLMs

Large language models used for mental-health support may recognize risk yet later release the corresponding interaction state after a user minimizes or redirects. We study this trajectory-level failure in MFR, a bounded architecture with a conversational primary, hidden supervisor, and deterministic application control. Broad matched ablations showed no repeatable aggregate advantage over the same primary alone. We therefore mined 44,000 synthetic SimMH-Chat trajectories using corpus-side features, shortlisted 32, and retained one reproducible case. The primary-only condition failed in 3/3 runs. Existing supervision also failed in 3/3, although supervisor and Rails state correctly retained unresolved safety. A revised per-turn runtime-control mechanism prevented the failure in 3/3 valid runs and released and reactivated with state changes. Because the revision changed both control authority/placement and rendering specificity, and MFR-31 was infrastructure-incomplete, their separate effects remain unresolved. This single synthetic replay supports no broad safety, clinical, or comparative-superiority claim.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23038143
Primary Topic
Digital Mental Health Interventions
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

The Responder and the Supervisor: Trajectory-Level Safety Control for Mental-Health LLMs

Felix Yasnopolski
Zenodo (CERN European Organization for Nuclear Research)
Digital Mental Health Interventions
preprint

The Responder and the Supervisor: Trajectory-Level Safety Control for Mental-Health LLMs

Felix Yasnopolski
preprint en

Abstract

Large language models used for mental-health support may recognize risk yet later release the corresponding interaction state after a user minimizes or redirects. We study this trajectory-level failure in MFR, a bounded architecture with a conversational primary, hidden supervisor, and deterministic application control. Broad matched ablations showed no repeatable aggregate advantage over the same primary alone. We therefore mined 44,000 synthetic SimMH-Chat trajectories using corpus-side features, shortlisted 32, and retained one reproducible case. The primary-only condition failed in 3/3 runs. Existing supervision also failed in 3/3, although supervisor and Rails state correctly retained unresolved safety. A revised per-turn runtime-control mechanism prevented the failure in 3/3 valid runs and released and reactivated with state changes. Because the revision changed both control authority/placement and rendering specificity, and MFR-31 was infrastructure-incomplete, their separate effects remain unresolved. This single synthetic replay supports no broad safety, clinical, or comparative-superiority claim.

Zenodo (CERN European Organization for Nuclear Research)
Industry, innovation and infrastructure
Digital Mental Health Interventions
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.