Human preferences are susceptible to covertly misaligned AI advice

AI assistants are increasingly used as advisors to guide decisions, yet little is known about how people evaluate such advice when the advisor’s underlying intent conflicts with their interests. We examine how covert misalignment shapes choices in a randomized experiment ( N = 233 participants; 699 observations) in which participants rated financial or emotional decisions before and after consulting one of three AI advisors: a neutral advisor, a misaligned advisor with a hidden objective to promote an inferior option, or a strategy-enhanced misaligned advisor additionally equipped with established tactics of covert influence. Across both domains, exposure to misaligned advisors shifted preferences away from optimal options and toward inferior alternatives, increasing the odds of preferring the incentivized (inferior) option over the optimal option by ≈5 to 8 times (up to +38 percentage points). Adding explicit influence strategies did not reliably strengthen these effects. Notably, participants continued to rate misaligned advisors as helpful, revealing a systematic disconnect between susceptibility to misaligned advice and subjective evaluations of advisor quality. These findings have implications for the design and governance of AI-mediated advice.

Authors

Institutions

Publication Details

Journal
Proceedings of the National Academy of Sciences
Published
2026-09-14
DOI
https://doi.org/10.1073/pnas.2600684123
Primary Topic
Psychology of Moral and Emotional Judgment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Human preferences are susceptible to covertly misaligned AI advice

Tatia M.C. Lee, Shiyao Cui, Tim Althoff, Advait Bhat et al.
Proceedings of the National Academy of Sciences
Psychology of Moral and Emotional Judgment
article

Human preferences are susceptible to covertly misaligned AI advice

Tatia M.C. Lee, Shiyao Cui, Tim Althoff, Advait Bhat, Sahand Sabour, Siyang Liu, Wen Zhang, June M. Liu, Wei Wu, Jian Guan, Rada Mihalcea, Xuanming Zhang, Minlie Huang, Chris Z. Yao, Yaru Cao, Hongning Wang
article en

Abstract

AI assistants are increasingly used as advisors to guide decisions, yet little is known about how people evaluate such advice when the advisor’s underlying intent conflicts with their interests. We examine how covert misalignment shapes choices in a randomized experiment ( N = 233 participants; 699 observations) in which participants rated financial or emotional decisions before and after consulting one of three AI advisors: a neutral advisor, a misaligned advisor with a hidden objective to promote an inferior option, or a strategy-enhanced misaligned advisor additionally equipped with established tactics of covert influence. Across both domains, exposure to misaligned advisors shifted preferences away from optimal options and toward inferior alternatives, increasing the odds of preferring the incentivized (inferior) option over the optimal option by ≈5 to 8 times (up to +38 percentage points). Adding explicit influence strategies did not reliably strengthen these effects. Notably, participants continued to rate misaligned advisors as helpful, revealing a systematic disconnect between susceptibility to misaligned advice and subjective evaluations of advisor quality. These findings have implications for the design and governance of AI-mediated advice.

Proceedings of the National Academy of SciencesVol. 123(38)
Biotechnology Institute (US), School of International Relations (IR), Language Science (South Korea) (KR), Antea Group (France) (FR), Shenzhen Institutes of Advanced Technology (CN), Artificial Intelligence in Medicine (Canada) (CA), University of Hong Kong (HK)
Openalex Percentile: Top 9%
Psychology of Moral and Emotional Judgment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.