The Sincerity Trap: Cascading Preference Falsification in Technology Risk Communication

Abstract Centralized-power institutions shaping artificial intelligence development produce public communication about risk that systematically understates what their senior actors privately assess, under the incentive conditions that Section 7 specifies. I argue that this gap is produced by cascading preference falsification across nested social layers — concentric circles of disclosure extending from private thought through professional in-groups to formal public output and downstream media — where each transition adjusts information toward the transmitter’s dominant prior. In the US political-economic context of frontier AI corporations and executive-branch AI policy, asymmetric reputational and financial incentives at each transmission point produce priors favoring reassurance, yielding a domain-scoped directional bias rather than a universal regularity. The mechanism is made invisible by what I call the Sincerity Trap: because no deception occurs at any individual layer, observers pattern-match the sincerity of each transmitter to the fidelity of the aggregate signal. Sincerity at each layer is not evidence of fidelity across layers. I document the mechanism through three cases of self-narrated filtering: Dean Ball, former primary author of the Trump Administration AI Action Plan; Dario Amodei, CEO of Anthropic; and Eric Schmidt, former CEO of Google and chair of the National Security Commission on AI. The framework synthesizes five core traditions (preference falsification, social penetration theory, dramaturgical sociology, social amplification of risk, and organizational silence), extended by two essential supporting traditions: exit-voice-loyalty dynamics, which explains the documented exit-to-voice pattern among former insiders, and scientific reticence, the closest prior analog in climate science. I propose five falsifiable predictions and argue that the resulting distortion degrades collective coordination capacity at the civilizational scale at which AI governance decisions are now being made. Working paper v1.11 (evidence through 16 September 2026). Supplement S1, the registered case-level disconfirmation search, is included in this record.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-17
DOI
https://doi.org/10.5281/zenodo.22814774
Primary Topic
Ethics and Social Impacts of AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

The Sincerity Trap: Cascading Preference Falsification in Technology Risk Communication

Turquoise Sound
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
article

The Sincerity Trap: Cascading Preference Falsification in Technology Risk Communication

Turquoise Sound
article en

Abstract

Abstract Centralized-power institutions shaping artificial intelligence development produce public communication about risk that systematically understates what their senior actors privately assess, under the incentive conditions that Section 7 specifies. I argue that this gap is produced by cascading preference falsification across nested social layers — concentric circles of disclosure extending from private thought through professional in-groups to formal public output and downstream media — where each transition adjusts information toward the transmitter’s dominant prior. In the US political-economic context of frontier AI corporations and executive-branch AI policy, asymmetric reputational and financial incentives at each transmission point produce priors favoring reassurance, yielding a domain-scoped directional bias rather than a universal regularity. The mechanism is made invisible by what I call the Sincerity Trap: because no deception occurs at any individual layer, observers pattern-match the sincerity of each transmitter to the fidelity of the aggregate signal. Sincerity at each layer is not evidence of fidelity across layers. I document the mechanism through three cases of self-narrated filtering: Dean Ball, former primary author of the Trump Administration AI Action Plan; Dario Amodei, CEO of Anthropic; and Eric Schmidt, former CEO of Google and chair of the National Security Commission on AI. The framework synthesizes five core traditions (preference falsification, social penetration theory, dramaturgical sociology, social amplification of risk, and organizational silence), extended by two essential supporting traditions: exit-voice-loyalty dynamics, which explains the documented exit-to-voice pattern among former insiders, and scientific reticence, the closest prior analog in climate science. I propose five falsifiable predictions and argue that the resulting distortion degrades collective coordination capacity at the civilizational scale at which AI governance decisions are now being made. Working paper v1.11 (evidence through 16 September 2026). Supplement S1, the registered case-level disconfirmation search, is included in this record.

Zenodo (CERN European Organization for Nuclear Research)
World Islamic Sciences and Education University (JO)
Climate action
Openalex Percentile: Top 7%
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.