Knows Its Slips, Reads the Page: A Pre-registered Study of Confidence and Abstention in Jev, a Calibrated Decision Model

Jev (TypeSafe AI, jev-1.13, released 15 September 2026) is a decision model. It does not write text. Given a situation and a list of options, it returns a choice and a probability for every option, and it is trained and sold for calibrated decisions. We treated it as comparative psychology treats animals that cannot report, and put two pre-registered questions to it in 48,300 calls costing US$1.53. Every stimulus was generated by code, with a known answer and an exactly computable ideal observer. J1 asked whether Jev's confidence knows its own competence. Its confidence predicted its own errors beyond a flexible model of the difficulty on the page (ΔAUROC +0.039 and +0.005 in two task families, both p = .0005). That is a measurable signal about its own processing. But when the same evidence was laid out less readably, its confidence fell, and it stayed low when stronger evidence brought accuracy back. On clean pages Jev was overconfident by 0.10; on reworked pages at matched accuracy it was underconfident by 0.04 (difference −0.136, 95% CI −0.173 to −0.099). Its confidence reads the presentation as well as itself. With every evidence value removed, it still chose at 0.82 mean confidence while performing at chance. J2 asked whether its decisions let go as content is removed. Given an explicit no-action option and a payoff table, Jev almost never held when acting was clearly optimal (0.02), almost always held when nothing was left (0.98), and held more when errors cost more (p = .0001). Its threshold for acting, though, barely moved with the stakes: 0.39–0.44 where the ideal is 0.50–0.90. Renaming the no-action option "witness: observe without acting" raised its probability by up to 0.28 in otherwise identical situations. A system's own vocabulary can move its decisions. That lesson was first measured in a language model's output (Vanhorn, 2026), and it returns here in a system that has no words of its own. The protocol, instrument, raw responses and analysis are public, and the registration was pushed before the first call. The paper opens with a plain-language account of what was learned and what it means for people building with Jev. The protocol, instrument, test suites, every raw response and the analysis are in the public repository linked below.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-26
DOI
https://doi.org/10.5281/zenodo.22980293
Primary Topic
Death Anxiety and Social Exclusion
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Knows Its Slips, Reads the Page: A Pre-registered Study of Confidence and Abstention in Jev, a Calibrated Decision Model

Joseph Vanhorn
Zenodo (CERN European Organization for Nuclear Research)
Death Anxiety and Social Exclusion
preprint

Knows Its Slips, Reads the Page: A Pre-registered Study of Confidence and Abstention in Jev, a Calibrated Decision Model

Joseph Vanhorn
preprint en

Abstract

Jev (TypeSafe AI, jev-1.13, released 15 September 2026) is a decision model. It does not write text. Given a situation and a list of options, it returns a choice and a probability for every option, and it is trained and sold for calibrated decisions. We treated it as comparative psychology treats animals that cannot report, and put two pre-registered questions to it in 48,300 calls costing US$1.53. Every stimulus was generated by code, with a known answer and an exactly computable ideal observer. J1 asked whether Jev's confidence knows its own competence. Its confidence predicted its own errors beyond a flexible model of the difficulty on the page (ΔAUROC +0.039 and +0.005 in two task families, both p = .0005). That is a measurable signal about its own processing. But when the same evidence was laid out less readably, its confidence fell, and it stayed low when stronger evidence brought accuracy back. On clean pages Jev was overconfident by 0.10; on reworked pages at matched accuracy it was underconfident by 0.04 (difference −0.136, 95% CI −0.173 to −0.099). Its confidence reads the presentation as well as itself. With every evidence value removed, it still chose at 0.82 mean confidence while performing at chance. J2 asked whether its decisions let go as content is removed. Given an explicit no-action option and a payoff table, Jev almost never held when acting was clearly optimal (0.02), almost always held when nothing was left (0.98), and held more when errors cost more (p = .0001). Its threshold for acting, though, barely moved with the stakes: 0.39–0.44 where the ideal is 0.50–0.90. Renaming the no-action option "witness: observe without acting" raised its probability by up to 0.28 in otherwise identical situations. A system's own vocabulary can move its decisions. That lesson was first measured in a language model's output (Vanhorn, 2026), and it returns here in a system that has no words of its own. The protocol, instrument, raw responses and analysis are public, and the registration was pushed before the first call. The paper opens with a plain-language account of what was learned and what it means for people building with Jev. The protocol, instrument, test suites, every raw response and the analysis are in the public repository linked below.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Death Anxiety and Social Exclusion
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Knows Its Slips, Reads the Page: A Pre-registered Study of Confidence and Abstention in Jev, a Calibrated Decision Model — Joseph Vanhorn · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS