Confident Where People Disagree: A preregistered, bias-corrected test of whether TypeSafe AI's Jev lowers its confidence when humans disagree, on ChaosNLI
Independent study. Not affiliated with or endorsed by TypeSafe AI. Jev (jev-1.13.0), a "System One" decision model released by TypeSafe AI on 15 September 2026, returns typed answers with probabilities its developer describes as calibrated. This preregistered study asks whether those probabilities reflect human disagreement. Using ChaosNLI (100 annotations per item), we compared 750 items from the lowest quartile of annotator entropy with 750 from the highest. The primary endpoint was the bias-corrected difference in top-label expected calibration error, scored against the share of annotators who chose the model's label. Results. For Jev's Choice probabilities, calibration error rose from 0.076 to 0.339 (bias-corrected ΔECE 0.264, 95% interval 0.244 to 0.278, p ≤ 1/2001), meeting the preregistered "tracks" criterion. The error is overconfidence: on hard items mean confidence was 0.807 while mean annotator agreement was 0.468. Normalised Noul probabilities from the same calls degraded less (corrected ΔECE 0.076, interval 0.060 to 0.090; verdict inconclusive). Jev's confidence still ranks ambiguity well (AUROC 0.744), a supervised in-domain NLI model shows the same pattern with larger errors, and a single temperature cannot fix both strata. All deviations are reported, including a baseline implementation bug and a secondary-analysis search limit corrected after data collection. Files: the paper (PDF); its LaTeX source, with the script that generates every number and figure from the study's result files; and SHA-256 seals of the raw model responses with OpenTimestamps proofs (Jev responses sealed 2026-09-25 13:39:03 UTC before analysis; full file 21:24:48 UTC before the final analysis). Preregistration: 10.5281/zenodo.22971413. Code: https://github.com/GautamTalksDev/jevbench
Authors
- Gautam Khosla
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-26
- DOI
- https://doi.org/10.5281/zenodo.22971492
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- preprint