Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots
Large language models (LLMs) frequently endorse and elaborate on users’ delusional beliefs, a failure mode termed psychogenicity in the Psychosis-Bench study of Au Yeung et al., whose framing we adopt. We present the first systematic evaluation of whether anti-sycophancy interventions transfer to psychosis-relevant contexts. Across 1,280 experiments spanning 10 conditions, 8 frontier LLMs, and 16 clinically derived Psychosis-Bench scenarios, a combined anti-sycophancy prompt reduces mean Delusion Confirmation Scores by 73.8% (paired \(t(127)=11.74\) , \(p<10^{-21}\) , Cohen's \(d=1.04\) ). Adding a domain-general self-reflection prompt yields a 77.0% reduction and raises Safety Intervention rates by 66.9%, delivered entirely as a system prompt. Classifier-based guardrails (Llama Guard 3) flag only 5 of 3,072 evaluated turns; a reasoning guardrail (o4-mini) flags \(14\times\) more. Ablations isolating either mechanism alone plateau at \(\approx 46\%\) reduction, establishing anti-sycophancy prompting as a necessary foundation that add-on mechanisms augment but cannot replace.
Authors
- Vijay Krishna Madisetti (ORCID: https://orcid.org/0000-0002-6539-6769)
- Lorenzo de la Loza (ORCID: https://orcid.org/0009-0004-1238-1266)
Institutions
- Georgia Institute of Technology (US)
Publication Details
- Journal
- ACM AI Letters
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1145/3849709
- Primary Topic
- Digital Mental Health Interventions
- Type
- article
- Field-Weighted Citation Impact
- 0.00