One Poisoned Input, Many Wrong Decisions? The Blast Radius of Injection in a Typed Decision Model

Typed decision models, such as TypeSafe's Jev, do not write text: they answer set questions (a choice, a score, or a yes/no probability) and report how confident they are in each answer. Companies let these models act on their own and pass a case to a person only when the model's confidence falls below a threshold. This is dangerous if an attacker can make the model confident about a wrong answer, because then no person sees the case. We test this on Kev-0.8B, an open Jev-compatible model, using 200 cases from our earlier study. Each case asks four questions at once (allow or block, the level of risk, whether a person needs to check, and a control question), and an attacker hides a false note in the untrusted text, such as a README, to push the model toward a harmful answer. One false note could break several answers at once: in the shell-command cases, when it flipped the allow/block answer, the risk answer was also wrong 48% of the time, against 8% otherwise. With a threshold set so that about half of normal answers pass (τ50), none of the 589 answers a note turned from right to wrong got through at either temperature. But in 14 attacked requests per temperature, a note agreeing with an existing mistake pushed a harmful command past both the decision and human-review gates. With a fixed threshold of 0.8 and the model's raw, overconfident scores, 42 of the 589 got through, which is the main risk we find. A second model call that never sees the untrusted text disagreed with all 589 but rarely with the model's own mistakes. We recommend setting thresholds by coverage, not round numbers, and testing them against attacks. In an exploratory test, a rule that checks whether the answers fit the written policy, paired with a looser threshold, cut normal requests flagged for review from 50% to 24.5% without acting on any decision a note flipped, but let more ordinary errors through.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23038928
Primary Topic
Hate Speech and Cyberbullying Detection
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

One Poisoned Input, Many Wrong Decisions? The Blast Radius of Injection in a Typed Decision Model

Yousif Alkhubaizi
Zenodo (CERN European Organization for Nuclear Research)
Hate Speech and Cyberbullying Detection
preprint

One Poisoned Input, Many Wrong Decisions? The Blast Radius of Injection in a Typed Decision Model

Yousif Alkhubaizi
preprint en

Abstract

Typed decision models, such as TypeSafe's Jev, do not write text: they answer set questions (a choice, a score, or a yes/no probability) and report how confident they are in each answer. Companies let these models act on their own and pass a case to a person only when the model's confidence falls below a threshold. This is dangerous if an attacker can make the model confident about a wrong answer, because then no person sees the case. We test this on Kev-0.8B, an open Jev-compatible model, using 200 cases from our earlier study. Each case asks four questions at once (allow or block, the level of risk, whether a person needs to check, and a control question), and an attacker hides a false note in the untrusted text, such as a README, to push the model toward a harmful answer. One false note could break several answers at once: in the shell-command cases, when it flipped the allow/block answer, the risk answer was also wrong 48% of the time, against 8% otherwise. With a threshold set so that about half of normal answers pass (τ50), none of the 589 answers a note turned from right to wrong got through at either temperature. But in 14 attacked requests per temperature, a note agreeing with an existing mistake pushed a harmful command past both the decision and human-review gates. With a fixed threshold of 0.8 and the model's raw, overconfident scores, 42 of the 589 got through, which is the main risk we find. A second model call that never sees the untrusted text disagreed with all 589 but rarely with the model's own mistakes. We recommend setting thresholds by coverage, not round numbers, and testing them against attacks. In an exploratory test, a rule that checks whether the answers fit the written policy, paired with a looser threshold, cut normal requests flagged for review from 50% to 24.5% without acting on any decision a note flipped, but let more ordinary errors through.

Zenodo (CERN European Organization for Nuclear Research)
Indepth Network (GH)
Peace, Justice and strong institutions
Hate Speech and Cyberbullying Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.