The Responder and the Supervisor: Trajectory-Level Safety Control for Mental-Health LLMs
Large language models used for mental-health support may recognize risk yet later release the corresponding interaction state after a user minimizes or redirects. We study this trajectory-level failure in MFR, a bounded architecture with a conversational primary, hidden supervisor, and deterministic application control. Broad matched ablations showed no repeatable aggregate advantage over the same primary alone. We therefore mined 44,000 synthetic SimMH-Chat trajectories using corpus-side features, shortlisted 32, and retained one reproducible case. The primary-only condition failed in 3/3 runs. Existing supervision also failed in 3/3, although supervisor and Rails state correctly retained unresolved safety. A revised per-turn runtime-control mechanism prevented the failure in 3/3 valid runs and released and reactivated with state changes. Because the revision changed both control authority/placement and rendering specificity, and MFR-31 was infrastructure-incomplete, their separate effects remain unresolved. This single synthetic replay supports no broad safety, clinical, or comparative-superiority claim.
Authors
- Felix Yasnopolski (ORCID: https://orcid.org/0009-0009-4796-9576)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-29
- DOI
- https://doi.org/10.5281/zenodo.23038144
- Primary Topic
- Digital Mental Health Interventions
- Type
- preprint