acceptodds
Under review as a conference paper at ICLR 2027

Characterizing the Separability Gap in Prior-Data Fitted Networks

Abstract

Prior-data fitted networks (PFNs) are transformers pre-trained on simulated datasets to approximate the posterior predictive distribution (PPD) of a task prior. Evaluating this approximation outside conjugate settings has remained challenging: for the priors of current PFNs the posterior is intractable, and real-world datasets lack ground-truth PPDs. A second difficulty is conceptual, as the standard formalization of PFNs assumes a separable prior, under which the feature distribution and the labeling mechanism are a priori independent. The structural-causal-model priors of current PFNs are non-separable, since features and labels are outputs of one shared random mechanism. We show that this changes the minimizer of the training objective and can invalidate the fixed-query martingale tests used to diagnose Bayesian behavior in PFNs. We characterize the resulting separability gap, show that it affects interpretability and Bayesian optimization, and exploit it to turn PFNs into semi-supervised predictors that use additional unlabeled rows. We characterize the gap of TabICL architectures pre-trained on these priors, test which predictive standard pre-training recovers, and provide insights into how PFNs behave as Bayesian predictors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.