Information Effects Are Not Incoherence: Fixed-Evidence Coherence Tests for Tabular Foundation Models

Tabular foundation models answer every query about a table by in-context inference, so whether their answers to related queries fit one joint distribution is a direct test of whether they can be used as probabilistic models of the table. A recent test of marginalization consistency compares a target’s marginal predicted with another column deleted from the training table against a mixture of conditionals predicted from tables that keep it, and finds violations in every model evaluated. We show that this comparison also changes the training evidence, and that a coherent Bayes predictor violates it: for a Bayes predictor the gap equals the information that the deleted column carries, up to the information of the second deleted column, it is exactly zero when the parameters of the two columns are a priori independent in a suitable factorization, and in a two-variable example it equals 12/25. On 20 synthetic families with exact Bayes posteriors and 8 to 128 training rows, the test flags the exact Bayes predictor in 14 families, while an evidence-fixed residual, which compares the two conditionals a model returns for the same table, stays at float64 round-off. Applied to TabICLv2 and TabPFNv2, the evidence-fixed residual is on average 0.19 and 0.25 of the column-deletion statistic, a descriptive ratio chosen after seeing the data. Measured against the change of the same two conditional predictions under a new ensemble seed, the residual exceeds three times this seed noise in 0 of 60 settings for either model (at most 6 when the noise is measured on other queries), while the column-deletion statistic does so in 34 and 35. At fixed training evidence, the incoherence of these two models is comparable to their ensemble-seed noise, and coherence tests for in-context predictors should hold the training evidence fixed.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23015075
Primary Topic
Bayesian Modeling and Causal Inference
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Information Effects Are Not Incoherence: Fixed-Evidence Coherence Tests for Tabular Foundation Models

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Bayesian Modeling and Causal Inference
preprint

Information Effects Are Not Incoherence: Fixed-Evidence Coherence Tests for Tabular Foundation Models

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

Tabular foundation models answer every query about a table by in-context inference, so whether their answers to related queries fit one joint distribution is a direct test of whether they can be used as probabilistic models of the table. A recent test of marginalization consistency compares a target’s marginal predicted with another column deleted from the training table against a mixture of conditionals predicted from tables that keep it, and finds violations in every model evaluated. We show that this comparison also changes the training evidence, and that a coherent Bayes predictor violates it: for a Bayes predictor the gap equals the information that the deleted column carries, up to the information of the second deleted column, it is exactly zero when the parameters of the two columns are a priori independent in a suitable factorization, and in a two-variable example it equals 12/25. On 20 synthetic families with exact Bayes posteriors and 8 to 128 training rows, the test flags the exact Bayes predictor in 14 families, while an evidence-fixed residual, which compares the two conditionals a model returns for the same table, stays at float64 round-off. Applied to TabICLv2 and TabPFNv2, the evidence-fixed residual is on average 0.19 and 0.25 of the column-deletion statistic, a descriptive ratio chosen after seeing the data. Measured against the change of the same two conditional predictions under a new ensemble seed, the residual exceeds three times this seed noise in 0 of 60 settings for either model (at most 6 when the noise is measured on other queries), while the column-deletion statistic does so in 34 and 35. At fixed training evidence, the incoherence of these two models is comparable to their ensemble-seed noise, and coherence tests for in-context predictors should hold the training evidence fixed.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW)
Reduced inequalities
Bayesian Modeling and Causal Inference
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.