Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots

Large language models (LLMs) frequently endorse and elaborate on users’ delusional beliefs, a failure mode termed psychogenicity in the Psychosis-Bench study of Au Yeung et al., whose framing we adopt. We present the first systematic evaluation of whether anti-sycophancy interventions transfer to psychosis-relevant contexts. Across 1,280 experiments spanning 10 conditions, 8 frontier LLMs, and 16 clinically derived Psychosis-Bench scenarios, a combined anti-sycophancy prompt reduces mean Delusion Confirmation Scores by 73.8% (paired \(t(127)=11.74\) , \(p<10^{-21}\) , Cohen's \(d=1.04\) ). Adding a domain-general self-reflection prompt yields a 77.0% reduction and raises Safety Intervention rates by 66.9%, delivered entirely as a system prompt. Classifier-based guardrails (Llama Guard 3) flag only 5 of 3,072 evaluated turns; a reasoning guardrail (o4-mini) flags \(14\times\) more. Ablations isolating either mechanism alone plateau at \(\approx 46\%\) reduction, establishing anti-sycophancy prompting as a necessary foundation that add-on mechanisms augment but cannot replace.

Authors

Institutions

Publication Details

Journal
ACM AI Letters
Published
2026-09-30
DOI
https://doi.org/10.1145/3849709
Primary Topic
Digital Mental Health Interventions
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots

Vijay Krishna Madisetti, Lorenzo de la Loza
ACM AI Letters
Digital Mental Health Interventions
article

Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots

Vijay Krishna Madisetti, Lorenzo de la Loza
article en

Abstract

Large language models (LLMs) frequently endorse and elaborate on users’ delusional beliefs, a failure mode termed psychogenicity in the Psychosis-Bench study of Au Yeung et al., whose framing we adopt. We present the first systematic evaluation of whether anti-sycophancy interventions transfer to psychosis-relevant contexts. Across 1,280 experiments spanning 10 conditions, 8 frontier LLMs, and 16 clinically derived Psychosis-Bench scenarios, a combined anti-sycophancy prompt reduces mean Delusion Confirmation Scores by 73.8% (paired \(t(127)=11.74\) , \(p<10^{-21}\) , Cohen's \(d=1.04\) ). Adding a domain-general self-reflection prompt yields a 77.0% reduction and raises Safety Intervention rates by 66.9%, delivered entirely as a system prompt. Classifier-based guardrails (Llama Guard 3) flag only 5 of 3,072 evaluated turns; a reasoning guardrail (o4-mini) flags \(14\times\) more. Ablations isolating either mechanism alone plateau at \(\approx 46\%\) reduction, establishing anti-sycophancy prompting as a necessary foundation that add-on mechanisms augment but cannot replace.

ACM AI Letters
Georgia Institute of Technology (US)
Openalex Percentile: Top 10%
Digital Mental Health Interventions
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots — Vijay Krishna Madisetti, Lorenzo de la Loza · ACM AI Letters (2026) | TGRS Research Map | TGRS