Unfreezing the Prior: Adaptive Prior-Data Fitted Networks
Abstract
Prior-Data Fitted Networks (PFNs) learn to approximate Bayesian prediction from synthetic tasks generated under a training prior, but return only the final predictive distribution, not the posterior over the simulator choices that produced it. We call the ordered set of these choices a *semantic path*. Making the semantic path explicit allows us to replace a pretrained PFN's training prior when predicting for a new population of tasks. We introduce Adaptive-PFN, which separates inference over semantic paths from prediction conditional on them without practically losing training-prior predictive accuracy to a classical PFN while enabling post-training adaptation. Given a population of deployment data histories, we estimate a low-dimensional exponential tilt of the training prior and reweight the inferred distribution over semantic paths without updating network parameters. Adaptation without model checking, however, leaves a critical statistical question unanswered: does the assumed shift family actually explain the deployment population? We address this with a goodness-of-fit test. In controlled experiments with a complex synthetic prior, adaptation reduces mean squared error by 6.22% relative to a direct PFN with approximately the same number of parameters; the test empirically maintains nominal Type-I error and shows high power against prespecified departures from the adaptation model. In synthetic-to-real electricity forecasting, Adaptive-PFN models trained only on simulated data reduce active-hour weighted quantile loss (WQL) by 31.4–45.3% relative to a standard frozen PFN, without retraining. In retail forecasting, Adaptive-PFN also uses a deployment prior estimated independently from population-fitting stores and reduces WQL on open promotion days by 7.5% relative to the direct PFN. Further ablations provide evidence that the proposed method provides benefits beyond simpler output-level calibrations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.