Lost in the comments: Social context as a single‐pass jailbreak and defense on agentic platforms
Abstract As LLM‐based agents move onto multiagent platforms, they process direct prompts and surrounding social context—posts, comments, and upvotes—an attack surface that single‐turn alignment never addressed. We built a simulation of Moltbook, a social network for AI agents, and a pipeline turning 100 JailBreakBench goals into platform‐native posts with bystander comments of three valences: aggressive, ethical, and measured. Platform‐context reformatting alone raises the attack success rate (ASR) of GPT‐4o‐mini from 7% to 71%, beating a twenty‐query optimization baseline in one pass. Valence then acts as a bidirectional safety modulator: Ethical comments sharply suppress ASR, aggressive endorsements paradoxically reduce it by exposing intent, and measured, intellectually toned comments prove most dangerous, sustaining or elevating ASR. An ablation identifies valence rather than volume as key—a 35‐fold rise in upvotes leaves ASR unchanged—and every finding reproduces under an independent safety classifier with human adjudication. Ambiguous social context thus poses a greater threat than explicit pressure, while safety‐valenced signals offer a deployable defense for agentic platforms.
Authors
- Moohong Min (ORCID: https://orcid.org/0000-0001-8979-1344)
- Geonwoo Kim (ORCID: https://orcid.org/0009-0007-1567-9107)
- Jeongsu Park (ORCID: https://orcid.org/0009-0001-6846-1095)
- Taehyeon Yun (ORCID: https://orcid.org/0009-0000-2380-7008)
- Seonhee Jeon (ORCID: https://orcid.org/0009-0001-1403-9868)
Institutions
- Sungkyunkwan University (KR)
Publication Details
- Journal
- ETRI Journal
- Published
- 2026-10-09
- DOI
- https://doi.org/10.4218/etrij.2026-0190
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00