Beyond Keywords: Attack Signatures and a Diagnostic Benchmark for Multilingual Scam Detection in Indian Languages

Scam-message classifiers are typically evaluated on held-out samples drawn from the same distribution as their training data, where high scores are easy to obtain and reveal little about whether a model has learned the structure of an attack or merely its vocabulary. We argue that the right unit of analysis is the Attack Signature: a typed tuple recording who the sender is impersonating, what pretext is deployed, which asset is sought, what the recipient is asked to do, and which persuasion levers are pulled. A signature should be invariant to surface realisation—the same fraud written in English, Hindi, Telugu, Tamil, Kannada, romanised transliteration or code-mixed script is one attack—and sensitive to security semantics, so that inverting a single clause (“share your OTP” → “never share your OTP”) changes it. We formalise this notion, define four diagnostic measures over it, and release Bharat-Scam-X, a constructed probe corpus of 1,823 labelled messages spanning nine Indian language varieties, twelve attack families and five evaluation conditions, including 143 minimal-edit contrast pairs whose members share 55.2% of their tokens but carry opposite security meaning. Evaluating four baselines, we find the gap we predicted. A character n-gram classifier reaches 0.942 macro-F1 in-distribution but only 0.700 on semantic flips, and marks 50.7% of genuine security advisories as attacks. A structured model that predicts the signature and derives risk from it is substantially more flip-sensitive (0.559 vs. 0.175 flip-success) and never fires on our hard negatives, yet it recovers the correct signature core for only 5.6% of recombined attacks that use exclusively training-seen components. We conclude that in-distribution accuracy is close to uninformative for this task, and release the corpus, schema and harness to make the harder questions measurable.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.22867073
Primary Topic
Spam and Phishing Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Beyond Keywords: Attack Signatures and a Diagnostic Benchmark for Multilingual Scam Detection in Indian Languages

shreya p, Ram Nivas attanti Attanti
Zenodo (CERN European Organization for Nuclear Research)
Spam and Phishing Detection
article

Beyond Keywords: Attack Signatures and a Diagnostic Benchmark for Multilingual Scam Detection in Indian Languages

shreya p, Ram Nivas attanti Attanti
article en

Abstract

Scam-message classifiers are typically evaluated on held-out samples drawn from the same distribution as their training data, where high scores are easy to obtain and reveal little about whether a model has learned the structure of an attack or merely its vocabulary. We argue that the right unit of analysis is the Attack Signature: a typed tuple recording who the sender is impersonating, what pretext is deployed, which asset is sought, what the recipient is asked to do, and which persuasion levers are pulled. A signature should be invariant to surface realisation—the same fraud written in English, Hindi, Telugu, Tamil, Kannada, romanised transliteration or code-mixed script is one attack—and sensitive to security semantics, so that inverting a single clause (“share your OTP” → “never share your OTP”) changes it. We formalise this notion, define four diagnostic measures over it, and release Bharat-Scam-X, a constructed probe corpus of 1,823 labelled messages spanning nine Indian language varieties, twelve attack families and five evaluation conditions, including 143 minimal-edit contrast pairs whose members share 55.2% of their tokens but carry opposite security meaning. Evaluating four baselines, we find the gap we predicted. A character n-gram classifier reaches 0.942 macro-F1 in-distribution but only 0.700 on semantic flips, and marks 50.7% of genuine security advisories as attacks. A structured model that predicts the signature and derives risk from it is substantially more flip-sensitive (0.559 vs. 0.175 flip-success) and never fires on our hard negatives, yet it recovers the correct signature core for only 5.6% of recombined attacks that use exclusively training-seen components. We conclude that in-distribution accuracy is close to uninformative for this task, and release the corpus, schema and harness to make the harder questions measurable.

Zenodo (CERN European Organization for Nuclear Research)
Jain University (IN)
Peace, Justice and strong institutions
Openalex Percentile: Top 4%
Spam and Phishing Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Beyond Keywords: Attack Signatures and a Diagnostic Benchmark for Multilingual Scam Detection in Indian Languages — shreya p, Ram Nivas attanti Attanti · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS