LexObf-12: Measuring and Hardening the Obfuscation Robustness of a Production Lexical Content Filter

Lexical filters remain the first stage of content moderation on many platforms because they are cheap, deterministic, and explainable, yet users evade them with substituted characters, inserted spaces, and look-alike Unicode letters. We introduce LexObf-12, a benchmark that applies twelve obfuscation classes to a filter's own lexicon in four carrier sentences, and use it, with two public labeled corpora and a list of common English words, to evaluate the pre-publication text filter of Kibhi (kibhi.com), a production short-video and messaging platform. The evaluation exposed two defects in the deployed normalizer: trailing punctuation was read as character substitution, so any listed phrase followed by "!" passed, and repeated-letter collapsing erased one listed term. Look-alike Cyrillic letters and zero-width characters defeated it entirely. After correction and hardening, mean detection across the twelve classes rose from 61.9% to 99.3%, against 17.7% and 28.7% for token and substring baselines, at 65 µs per check and with no change in false positives on 24,783 labeled tweets (1.95%) or among 10,000 common words. A leave-one-out ablation attributes the gain to each step and shows that only plural matching affects false positives. On 19,229 HateXplain posts, however, the filter flagged 12.0% of posts that annotators labeled normal, mostly through slur matches, and about 90% of the hate posts it missed contained no listed term at all. We conclude that normalization should be tested adversarially, that the limit of a word list is coverage rather than matching, and that slur matches should be held for human review rather than blocked. The benchmark generator is released.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23147385
Primary Topic
Hate Speech and Cyberbullying Detection
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

LexObf-12: Measuring and Hardening the Obfuscation Robustness of a Production Lexical Content Filter

Tafadzwa Tauro
Zenodo (CERN European Organization for Nuclear Research)
Hate Speech and Cyberbullying Detection
preprint

LexObf-12: Measuring and Hardening the Obfuscation Robustness of a Production Lexical Content Filter

Tafadzwa Tauro
preprint en

Abstract

Lexical filters remain the first stage of content moderation on many platforms because they are cheap, deterministic, and explainable, yet users evade them with substituted characters, inserted spaces, and look-alike Unicode letters. We introduce LexObf-12, a benchmark that applies twelve obfuscation classes to a filter's own lexicon in four carrier sentences, and use it, with two public labeled corpora and a list of common English words, to evaluate the pre-publication text filter of Kibhi (kibhi.com), a production short-video and messaging platform. The evaluation exposed two defects in the deployed normalizer: trailing punctuation was read as character substitution, so any listed phrase followed by "!" passed, and repeated-letter collapsing erased one listed term. Look-alike Cyrillic letters and zero-width characters defeated it entirely. After correction and hardening, mean detection across the twelve classes rose from 61.9% to 99.3%, against 17.7% and 28.7% for token and substring baselines, at 65 µs per check and with no change in false positives on 24,783 labeled tweets (1.95%) or among 10,000 common words. A leave-one-out ablation attributes the gain to each step and shows that only plural matching affects false positives. On 19,229 HateXplain posts, however, the filter flagged 12.0% of posts that annotators labeled normal, mostly through slur matches, and about 90% of the hate posts it missed contained no listed term at all. We conclude that normalization should be tested adversarially, that the limit of a word list is coverage rather than matching, and that slur matches should be held for human review rather than blocked. The benchmark generator is released.

Zenodo (CERN European Organization for Nuclear Research)
Hate Speech and Cyberbullying Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.