Information-Set Confounding in Consistency Tests for Tabular Foundation Models
Abstract
Recent work tests compatibility of univariate tabular foundation model (TFM) predictions by comparing marginals with autoregressive conditionals. The tested identities, however, combine posterior predictives based on different observed training columns. Because the law of total probability and the chain rule apply under a common information set, a nonzero cross-information gap need not imply a failure of Bayesian coherence across observed datasets. After making the conditioning explicit, we bound both reported gaps by ordinary posterior updates. We also construct an exact coherent Bayesian model for which both gaps equal , while a predictor that ignores cross-target information attains zero gap at strictly worse joint risk. This reversal appears in a 60-task exact Bayesian benchmark, controlled experiments with TabICLv2 and TabPFNv2, and all ten classification datasets in the original study. On each dataset and for both models, a fixed mixture of the native conditional chains improves joint negative log likelihood over the zero-gap independent product; the combined pooled relative NLL gains are 10.62% and 10.17%. The queried heads can be incompatible when their different contexts are suppressed; our result shows why this does not establish that the predictor lacks a coherent Bayesian interpretation. A valid consistency test must hold the observed information fixed; practical evaluations should also include an information-blind control and held-out predictive scores.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.