Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
LLM agents are deployed behind dashboards, retrieval pipelines and market feeds on the assumption that more context makes their decisions more reliable. For questions that are irreducibly uncertain, that assumption inverts. Across 12 frontier models, commitment to a directional call on a provably unpredictable question rises from 6.5% to 54.0% as evidence is escalated. Fabricating the entire evidence panel, so that nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. It is not incapacity: on matched answerable questions attached to the same panels the same models answer essentially always, at ceiling accuracy. It is not belief: stated probabilities are anti-predictive of outcomes (AUROC 0.346), so an audit that reads the number cannot see the problem. It is not missing judgment: asked to classify a question's knowability first, models call it irreducible 90% of the time and then commit on 0.4% of those. The act/don't-act gate is what fails. Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases about dice and coins drives commitment to 0.0% on the original cases and transfers to three unseen domains across six independent runs. The boundary has a mechanism: the gate holds exactly when the response format leaves the model room to reason. Version 2 extends the earlier paper with three further domains, a fully fabricated evidence control, a bin-free calibration decomposition, a training intervention, and a map of where that intervention breaks. All cached model outputs, analysis code and the pre-registration are released.
Authors
- Pranav Aggarwal (ORCID: https://orcid.org/0009-0005-1243-0520)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22958012
- Primary Topic
- Library Science and Information Systems
- Type
- preprint