Knows Its Slips, Reads the Page: A Pre-registered Study of Confidence and Abstention in Jev, a Calibrated Decision Model
Jev (TypeSafe AI, jev-1.13, released 15 September 2026) is a decision model. It does not write text. Given a situation and a list of options, it returns a choice and a probability for every option, and it is trained and sold for calibrated decisions. We treated it as comparative psychology treats animals that cannot report, and put two pre-registered questions to it in 48,300 calls costing US$1.53. Every stimulus was generated by code, with a known answer and an exactly computable ideal observer. J1 asked whether Jev's confidence knows its own competence. Its confidence predicted its own errors beyond a flexible model of the difficulty on the page (ΔAUROC +0.039 and +0.005 in two task families, both p = .0005). That is a measurable signal about its own processing. But when the same evidence was laid out less readably, its confidence fell, and it stayed low when stronger evidence brought accuracy back. On clean pages Jev was overconfident by 0.10; on reworked pages at matched accuracy it was underconfident by 0.04 (difference −0.136, 95% CI −0.173 to −0.099). Its confidence reads the presentation as well as itself. With every evidence value removed, it still chose at 0.82 mean confidence while performing at chance. J2 asked whether its decisions let go as content is removed. Given an explicit no-action option and a payoff table, Jev almost never held when acting was clearly optimal (0.02), almost always held when nothing was left (0.98), and held more when errors cost more (p = .0001). Its threshold for acting, though, barely moved with the stakes: 0.39–0.44 where the ideal is 0.50–0.90. Renaming the no-action option "witness: observe without acting" raised its probability by up to 0.28 in otherwise identical situations. A system's own vocabulary can move its decisions. That lesson was first measured in a language model's output (Vanhorn, 2026), and it returns here in a system that has no words of its own. The protocol, instrument, raw responses and analysis are public, and the registration was pushed before the first call. The paper opens with a plain-language account of what was learned and what it means for people building with Jev. The protocol, instrument, test suites, every raw response and the analysis are in the public repository linked below.
Authors
- Joseph Vanhorn
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-26
- DOI
- https://doi.org/10.5281/zenodo.22980293
- Primary Topic
- Death Anxiety and Social Exclusion
- Type
- preprint