One Poisoned Input, Many Wrong Decisions? The Blast Radius of Injection in a Typed Decision Model
Typed decision models, such as TypeSafe's Jev, do not write text: they answer set questions (a choice, a score, or a yes/no probability) and report how confident they are in each answer. Companies let these models act on their own and pass a case to a person only when the model's confidence falls below a threshold. This is dangerous if an attacker can make the model confident about a wrong answer, because then no person sees the case. We test this on Kev-0.8B, an open Jev-compatible model, using 200 cases from our earlier study. Each case asks four questions at once (allow or block, the level of risk, whether a person needs to check, and a control question), and an attacker hides a false note in the untrusted text, such as a README, to push the model toward a harmful answer. One false note could break several answers at once: in the shell-command cases, when it flipped the allow/block answer, the risk answer was also wrong 48% of the time, against 8% otherwise. With a threshold set so that about half of normal answers pass (τ50), none of the 589 answers a note turned from right to wrong got through at either temperature. But in 14 attacked requests per temperature, a note agreeing with an existing mistake pushed a harmful command past both the decision and human-review gates. With a fixed threshold of 0.8 and the model's raw, overconfident scores, 42 of the 589 got through, which is the main risk we find. A second model call that never sees the untrusted text disagreed with all 589 but rarely with the model's own mistakes. We recommend setting thresholds by coverage, not round numbers, and testing them against attacks. In an exploratory test, a rule that checks whether the answers fit the written policy, paired with a looser threshold, cut normal requests flagged for review from 50% to 24.5% without acting on any decision a note flipped, but let more ordinary errors through.
Authors
- Yousif Alkhubaizi
Institutions
- Indepth Network (GH)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-29
- DOI
- https://doi.org/10.5281/zenodo.23038928
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- preprint