Separating Supervision Randomness in Prior-Data Fitted Networks
Abstract
Prior-data fitted networks (PFNs) learn predictive distributions from simulated tasks. At a fixed training budget, removing supervision randomness can improve source-prior fidelity while worsening prediction under prior shift, despite preserving the expected training objective. A four-arm intervention in a finite structural equation model separates additive component-target variation from within-component label randomness. Across five paired training seeds within one source prior, replacing posterior-resampled labels with component targets reduces source predictive KL by 0.00975 nats per query but increases coefficient-shift KL by 0.00413 at . A signed loss estimator retains within-component label randomness while removing additive component-target variation and reproduces the reversal. At the same interpolation level, including the residual instead increases predictive KL under covariance-family shift. A separate six-prior, two-arm study shows that the pooled supervision ranking changes with shift magnitude and construction: full-predictive supervision wins at the smaller tested shift and sampled supervision at the larger one in both families. Source diagnostics show that residuals dominate native target variance. Source-selected output softening remains competitive in aggregate risk while moving predictions farther from the residual-trained predictor on source assessment inputs. These findings establish supervision-estimator choice as a consequential variable in finite-budget PFN training, requiring joint evaluation of source fidelity, shifted risk and predictive behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.