acceptodds
Under review as a conference paper at ICLR 2027

Towards Principled Data Generation for Prior-Data Fitted Networks

Abstract

As natural sources of data are becoming exhausted, synthetic data is playing an increasingly important role for pre-training foundation models. Perhaps surprisingly, synthetic data has been shown to surpass natural data in downstream performance. For example, prior-data fitted networks (PFNs), which create an implicit prior through data generation, are trained entirely on synthetic data, but provide state-of-the-art results in tabular prediction. Despite this success, the design of synthetic data remains largely heuristic, requiring expensive trial-and-error rather than principled understanding. In this paper, we work towards a science of data generation, using PFNs to systematically understand the relationship between properties of the data and its performance in downstream tasks. We demonstrate that the synthetic data generation procedure shapes the training dynamics and inductive biases of the model, and thus the optimal pre-training data distribution may be distinct from the downstream task distribution. We identify two key properties of the prior, diversity and complexity, and show that increasing either consistently improves generalization, even if the prior is moved farther from the downstream task. To operationalize these ideas, we also introduce practical proxies and demonstrate that they are predictive of real-world downstream performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.