Lost in the comments: Social context as a single‐pass jailbreak and defense on agentic platforms

Abstract As LLM‐based agents move onto multiagent platforms, they process direct prompts and surrounding social context—posts, comments, and upvotes—an attack surface that single‐turn alignment never addressed. We built a simulation of Moltbook, a social network for AI agents, and a pipeline turning 100 JailBreakBench goals into platform‐native posts with bystander comments of three valences: aggressive, ethical, and measured. Platform‐context reformatting alone raises the attack success rate (ASR) of GPT‐4o‐mini from 7% to 71%, beating a twenty‐query optimization baseline in one pass. Valence then acts as a bidirectional safety modulator: Ethical comments sharply suppress ASR, aggressive endorsements paradoxically reduce it by exposing intent, and measured, intellectually toned comments prove most dangerous, sustaining or elevating ASR. An ablation identifies valence rather than volume as key—a 35‐fold rise in upvotes leaves ASR unchanged—and every finding reproduces under an independent safety classifier with human adjudication. Ambiguous social context thus poses a greater threat than explicit pressure, while safety‐valenced signals offer a deployable defense for agentic platforms.

Authors

Institutions

Publication Details

Journal
ETRI Journal
Published
2026-10-09
DOI
https://doi.org/10.4218/etrij.2026-0190
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Lost in the comments: Social context as a single‐pass jailbreak and defense on agentic platforms

Moohong Min, Geonwoo Kim, Jeongsu Park, Taehyeon Yun et al.
ETRI Journal
Adversarial Robustness in Machine Learning
article

Lost in the comments: Social context as a single‐pass jailbreak and defense on agentic platforms

Moohong Min, Geonwoo Kim, Jeongsu Park, Taehyeon Yun, Seonhee Jeon
article en

Abstract

Abstract As LLM‐based agents move onto multiagent platforms, they process direct prompts and surrounding social context—posts, comments, and upvotes—an attack surface that single‐turn alignment never addressed. We built a simulation of Moltbook, a social network for AI agents, and a pipeline turning 100 JailBreakBench goals into platform‐native posts with bystander comments of three valences: aggressive, ethical, and measured. Platform‐context reformatting alone raises the attack success rate (ASR) of GPT‐4o‐mini from 7% to 71%, beating a twenty‐query optimization baseline in one pass. Valence then acts as a bidirectional safety modulator: Ethical comments sharply suppress ASR, aggressive endorsements paradoxically reduce it by exposing intent, and measured, intellectually toned comments prove most dangerous, sustaining or elevating ASR. An ablation identifies valence rather than volume as key—a 35‐fold rise in upvotes leaves ASR unchanged—and every finding reproduces under an independent safety classifier with human adjudication. Ambiguous social context thus poses a greater threat than explicit pressure, while safety‐valenced signals offer a deployable defense for agentic platforms.

ETRI Journal
Sungkyunkwan University (KR)
Openalex Percentile: Top 12%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Lost in the comments: Social context as a single‐pass jailbreak and defense on agentic platforms — Moohong Min, Geonwoo Kim, et al. · ETRI Journal (2026) | TGRS Research Map | TGRS