Characterizing the Separability Gap in Prior-Data Fitted Networks
Abstract
Prior-data fitted networks (PFNs) are transformers pre-trained on simulated datasets to approximate the posterior predictive distribution (PPD) of a task prior. Evaluating this approximation outside conjugate settings has remained challenging: for the priors of current PFNs the posterior is intractable, and real-world datasets lack ground-truth PPDs. A second difficulty is conceptual, as the standard formalization of PFNs assumes a separable prior, under which the feature distribution and the labeling mechanism are a priori independent. The structural-causal-model priors of current PFNs are non-separable, since features and labels are outputs of one shared random mechanism. We show that this changes the minimizer of the training objective and can invalidate the fixed-query martingale tests used to diagnose Bayesian behavior in PFNs. We characterize the resulting separability gap, show that it affects interpretability and Bayesian optimization, and exploit it to turn PFNs into semi-supervised predictors that use additional unlabeled rows. We characterize the gap of TabICL architectures pre-trained on these priors, test which predictive standard pre-training recovers, and provide insights into how PFNs behave as Bayesian predictors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.