Evaluating Safety Embedding Prefiltering for Analyzing Millions of LLM Agent Social Network Messages for Security and Safety Harms
Agentic social networks expose autonomous agents to large volumes of potentially malicious content, including prompt injection, social engineering, and unsafe execution requests. Large Language Model (LLM) judges can detect such content, but applying them at platform scale is expensive. We study the extent to which lightweight embedding-based prefilters can reduce this cost while retaining most of the judge’s unsafe predictions. We annotate 10,000 Moltbook posts and comments with a frontier LLM judge, assigning a binary safety verdict, severity level, malicious intent taxonomy labels, and Open Worldwide Application Security Project (OWASP) AI risk codes. Of the LLM-annotated samples, 9.2% were unsafe at severity 3 or above. Two human annotators with high inter-rater agreement (Cohen’s κ=0.823) exhibited moderate agreement between their adjudicated labels and the LLM ratings (κ=0.578). We compare centroid-based cosine prefilters over 3 off-the-shelf encoders against a supervised trained classifier. Reducing the input window size raises average precision by up to 6.3 percentage points. At an operating point calibrated to 0.80 recall, the projected cost of scanning 787,226 messages falls from $7085 by 49.9% (MiniLM-L12-v2), 55.6% (BGE-M3), and 65.6% (trained classifier), while MiniLM runs roughly 32 times faster than the classifier. Prefiltering can therefore make frontier-judge screening affordable at scale, although jailbreak content remains the hardest category to retrieve.
Authors
- Ioana Brănescu
- Mihai Dascălu (ORCID: https://orcid.org/0000-0002-4815-9227)
- Traian Eugen Rebedea (ORCID: https://orcid.org/0000-0002-7255-5537)
- Paul-Ioan Clotan
- Gabriela Adelina Gherghe
Institutions
- Academia Oamenilor de Știință din România (RO)
- Nvidia (United States) (US)
- Universitatea Națională de Știință și Tehnologie Politehnica București (RO)
Publication Details
- Journal
- AI
- Published
- 2026-10-06
- DOI
- https://doi.org/10.3390/ai7100409
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00