Evaluating Safety Embedding Prefiltering for Analyzing Millions of LLM Agent Social Network Messages for Security and Safety Harms

Agentic social networks expose autonomous agents to large volumes of potentially malicious content, including prompt injection, social engineering, and unsafe execution requests. Large Language Model (LLM) judges can detect such content, but applying them at platform scale is expensive. We study the extent to which lightweight embedding-based prefilters can reduce this cost while retaining most of the judge’s unsafe predictions. We annotate 10,000 Moltbook posts and comments with a frontier LLM judge, assigning a binary safety verdict, severity level, malicious intent taxonomy labels, and Open Worldwide Application Security Project (OWASP) AI risk codes. Of the LLM-annotated samples, 9.2% were unsafe at severity 3 or above. Two human annotators with high inter-rater agreement (Cohen’s κ=0.823) exhibited moderate agreement between their adjudicated labels and the LLM ratings (κ=0.578). We compare centroid-based cosine prefilters over 3 off-the-shelf encoders against a supervised trained classifier. Reducing the input window size raises average precision by up to 6.3 percentage points. At an operating point calibrated to 0.80 recall, the projected cost of scanning 787,226 messages falls from $7085 by 49.9% (MiniLM-L12-v2), 55.6% (BGE-M3), and 65.6% (trained classifier), while MiniLM runs roughly 32 times faster than the classifier. Prefiltering can therefore make frontier-judge screening affordable at scale, although jailbreak content remains the hardest category to retrieve.

Authors

Institutions

Publication Details

Journal
AI
Published
2026-10-06
DOI
https://doi.org/10.3390/ai7100409
Primary Topic
Hate Speech and Cyberbullying Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Evaluating Safety Embedding Prefiltering for Analyzing Millions of LLM Agent Social Network Messages for Security and Safety Harms

Ioana Brănescu, Mihai Dascălu, Traian Eugen Rebedea, Paul-Ioan Clotan et al.
AI
Hate Speech and Cyberbullying Detection
article

Evaluating Safety Embedding Prefiltering for Analyzing Millions of LLM Agent Social Network Messages for Security and Safety Harms

Ioana Brănescu, Mihai Dascălu, Traian Eugen Rebedea, Paul-Ioan Clotan, Gabriela Adelina Gherghe
article en

Abstract

Agentic social networks expose autonomous agents to large volumes of potentially malicious content, including prompt injection, social engineering, and unsafe execution requests. Large Language Model (LLM) judges can detect such content, but applying them at platform scale is expensive. We study the extent to which lightweight embedding-based prefilters can reduce this cost while retaining most of the judge’s unsafe predictions. We annotate 10,000 Moltbook posts and comments with a frontier LLM judge, assigning a binary safety verdict, severity level, malicious intent taxonomy labels, and Open Worldwide Application Security Project (OWASP) AI risk codes. Of the LLM-annotated samples, 9.2% were unsafe at severity 3 or above. Two human annotators with high inter-rater agreement (Cohen’s κ=0.823) exhibited moderate agreement between their adjudicated labels and the LLM ratings (κ=0.578). We compare centroid-based cosine prefilters over 3 off-the-shelf encoders against a supervised trained classifier. Reducing the input window size raises average precision by up to 6.3 percentage points. At an operating point calibrated to 0.80 recall, the projected cost of scanning 787,226 messages falls from $7085 by 49.9% (MiniLM-L12-v2), 55.6% (BGE-M3), and 65.6% (trained classifier), while MiniLM runs roughly 32 times faster than the classifier. Prefiltering can therefore make frontier-judge screening affordable at scale, although jailbreak content remains the hardest category to retrieve.

AIVol. 7(10)
Academia Oamenilor de Știință din România (RO), Nvidia (United States) (US), Universitatea Națională de Știință și Tehnologie Politehnica București (RO)
Openalex Percentile: Top 11%
Hate Speech and Cyberbullying Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.