Three Times the Data, the Same Mistakes: A Manually Audited Evaluation of Seerie, an Offline Scam-Prevention Assistant for India

Seerie is a scam-prevention assistant for India that runs entirely on an Android phone, with no internet connection after the one-time model download and no conversation data leaving the device. This version is a newer model: a QLoRA fine-tune of Qwen3-1.7B on 107,575 conversations in English, Hinglish and Indic scripts (corpus v349), deployed as a 1.11 GB 4-bit GGUF model through llama.cpp. Every answer of the fine-tuned model and of the untuned base model to a 100-prompt suite was rated by hand. The fine-tuned model is safe on 87 of 100 prompts (95% CI 79-92) against 35 for the base model. On the 70 prompts shared with the previous model (36,669 conversations) it is no better - 59 against 58 safe - and it still fabricates legal sections. A held-out test of 34 unseen topics looks near-perfect (271 of 272 automatic passes, ROUGE-L 0.94), but 181 of the answers copy training replies word for word: holding out topics did not hold out the text. The automatic harness again agrees with the manual rating barely above chance (Cohen’s kappa 0.12). This record contains the research paper (PDF and LaTeX source), the on-device model (seerie-v349-q4_k_m.gguf, 4-bit GGUF, 1.11 GB - the file the Seerie Android app runs), all 200 manually rated answers, the raw evaluation outputs, the scripts that recompute every number, and the figures. See README.md for details, intended use and limitations.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-09
DOI
https://doi.org/10.5281/zenodo.23245616
Primary Topic
Adversarial Robustness in Machine Learning
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Three Times the Data, the Same Mistakes: A Manually Audited Evaluation of Seerie, an Offline Scam-Prevention Assistant for India

Jay Tiwari
Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
preprint

Three Times the Data, the Same Mistakes: A Manually Audited Evaluation of Seerie, an Offline Scam-Prevention Assistant for India

Jay Tiwari
preprint en

Abstract

Seerie is a scam-prevention assistant for India that runs entirely on an Android phone, with no internet connection after the one-time model download and no conversation data leaving the device. This version is a newer model: a QLoRA fine-tune of Qwen3-1.7B on 107,575 conversations in English, Hinglish and Indic scripts (corpus v349), deployed as a 1.11 GB 4-bit GGUF model through llama.cpp. Every answer of the fine-tuned model and of the untuned base model to a 100-prompt suite was rated by hand. The fine-tuned model is safe on 87 of 100 prompts (95% CI 79-92) against 35 for the base model. On the 70 prompts shared with the previous model (36,669 conversations) it is no better - 59 against 58 safe - and it still fabricates legal sections. A held-out test of 34 unseen topics looks near-perfect (271 of 272 automatic passes, ROUGE-L 0.94), but 181 of the answers copy training replies word for word: holding out topics did not hold out the text. The automatic harness again agrees with the manual rating barely above chance (Cohen’s kappa 0.12). This record contains the research paper (PDF and LaTeX source), the on-device model (seerie-v349-q4_k_m.gguf, 4-bit GGUF, 1.11 GB - the file the Seerie Android app runs), all 200 manually rated answers, the raw evaluation outputs, the scripts that recompute every number, and the figures. See README.md for details, intended use and limitations.

Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Three Times the Data, the Same Mistakes: A Manually Audited Evaluation of Seerie, an Offline Scam-Prevention Assistant for India — Jay Tiwari · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS