Look When Unsure, Check When Sure: Consequence Training Makes a World Model's Remaining Errors Confident, Most of All Where It Knows the World Best

An agent that predicts the consequences of its actions can chain those predictions and plan without acting, but errors compound and checking the real state costs time. With V42, the first release candidate of the open one-pass consequence model Ekbasis-27B, a simple rule (look when the chain's confidence falls below 0.9, plus Trickle-style scheduled checks) keeps 197 of 200 fresh long chains exact at 17.9 looks per 100 actions. The checks are needed because of a pre-registered finding: 58.1% of the model's errors in trained families carry confidence of at least 0.9, against 27.8% in unseen families, and the same questions answered by the model before consequence training show almost none. Recalibration and smoother losses do not fix it; training on errors mined inside the model's own chains does, at the cost of rare git knowledge; an exact weight interpolation halfway back to V42 keeps the gain without the cost, passes the release evaluation and is the released Ekbasis-27B. Every analysis is pre-registered and re-run from the saved outputs.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-04
DOI
https://doi.org/10.5281/zenodo.23146970
Primary Topic
AI-based Problem Solving and Planning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Look When Unsure, Check When Sure: Consequence Training Makes a World Model's Remaining Errors Confident, Most of All Where It Knows the World Best

Caio Vicentino
Zenodo (CERN European Organization for Nuclear Research)
AI-based Problem Solving and Planning
article

Look When Unsure, Check When Sure: Consequence Training Makes a World Model's Remaining Errors Confident, Most of All Where It Knows the World Best

Caio Vicentino
article en

Abstract

An agent that predicts the consequences of its actions can chain those predictions and plan without acting, but errors compound and checking the real state costs time. With V42, the first release candidate of the open one-pass consequence model Ekbasis-27B, a simple rule (look when the chain's confidence falls below 0.9, plus Trickle-style scheduled checks) keeps 197 of 200 fresh long chains exact at 17.9 looks per 100 actions. The checks are needed because of a pre-registered finding: 58.1% of the model's errors in trained families carry confidence of at least 0.9, against 27.8% in unseen families, and the same questions answered by the model before consequence training show almost none. Recalibration and smoother losses do not fix it; training on errors mined inside the model's own chains does, at the cost of rare git knowledge; an exact weight interpolation halfway back to V42 keeps the gain without the cost, passes the release evaluation and is the released Ekbasis-27B. Every analysis is pre-registered and re-run from the saved outputs.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 11%
AI-based Problem Solving and Planning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Look When Unsure, Check When Sure: Consequence Training Makes a World Model's Remaining Errors Confident, Most of All Where It Knows the World Best — Caio Vicentino · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS