acceptodds
Under review as a conference paper at ICLR 2027

STRUCTURE-ALIGNED SYNTHETIC PRIORS FOR TABULAR FOUNDATION MODELS

Abstract

Tabular foundation models such as TabPFN, TabICL, and MITRA are pre-trained on millions of synthetic datasets sampled from randomly generated Structural Causal Models (SCMs). In the dominant paradigm the SCM is realised as a randomly wired multi-layer perceptron (MLP) whose activations are read off as features and targets, so the dependency graph of the training data is an artifact of MLP connectivity rather than a deliberate choice. We argue that this topol- ogy is an independent design variable, and that it diverges systematically from real causal structure. Comparing∼36,000 TabICL SCMs against 33 real-world causal graphs of three independent provenances, we find that MLP priors produce conditional dependence structures an order of magnitude wider: median Markov- blanket width 43.8 versus 4.0, maximum parent count 31 versus 3, maximum minimum-d-separator cardinality 20 versus 3. This is not a restatement of graph size—MLP blanket width scales with node count (∝N0.73) whereas real-world width is size-invariant (N0.06, n.s.), and the gap persists at 10.2×within a matched size band. We therefore introduce BN-ICL, a prior whose dependency graph is an explicit Bayesian network matched to real structural statistics, and train in-context learners from scratch under an ablation in which the only difference is DAG topol- ogy. On an augmented suite of 219 real classification datasets (TALENT plus a high-dimensional OpenML expansion specified in advance), the BN prior im- proves mean accuracy and log-loss, but not uniformly: the effect is concentrated in, and grows with, feature dimensionality, and the priors are indistinguishable on the median dataset. Above 100 features the BN prior wins by +1.5–1.6 percent- age points (bootstrap 95% CI excluding zero, Holm-corrected p<0.05), replicat- ing across two independently trained prior configurations and on external datasets disjoint from TALENT—ruling out a benchmark-composition artifact—and sur- viving truncation of those tables to the training feature width when the retained features are informative, which rules out width extrapolation as the explanation. Low- and mid-dimensional gains are smaller or fragile. DAG topology is thus a meaningful, under-explored axis of prior design, sharpest where the structural discrepancy is largest.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.