acceptodds
Under review as a conference paper at ICLR 2027

PriorFuse: How Prior-Data Fitted Networks Depart from Bayes

Abstract

Prior-Data Fitted Networks (PFNs) let researchers specify prediction rules through synthetic task priors and are trained to approximate the corresponding Bayesian posterior predictive. However, specifying a Bayesian target does not guarantee that a finite trained network will use new evidence as that target would. To study this question, we introduce PriorFuse, an exactly enumerable benchmark in which the target posterior predictive can be computed exactly by Bayesian model averaging (BMA). For five PFNs trained only with eight context examples, supplying twelve examples improves predictive log loss while increasing excess log loss over exact BMA approximately ninefold at 100k training steps. We show that the trained PFNs exhibit a systematic Bayesian departure: fitted likelihood weighting changes from amplification at several shorter context lengths to attenuation at longer lengths. At the tested longer lengths, a post hoc inverse-length adjustment to the exponent fitted at eight examples recovers about 99% of the improvement in matching PFN predictions achieved by refitting at each length, and unchanged transfer recovers much less. The same matched-length attenuation pattern also appears under a second task prior and weakens with further training. Equivalent representations of the same prior also change fitted family parameters without changing either predictor. PriorFuse thus helps researchers distinguish better prediction from more faithful inference and identify which parts of a learned prediction rule change with training and context.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.