acceptodds
Under review as a conference paper at ICLR 2027

Fast In-Context Tabular Diffusion via Synthetic Pretraining

Abstract

Diffusion models for tabular data generation require training on each new dataset, making their use on new datasets computationally expensive. We introduce DiffPFN, an in-context diffusion model that generates tabular data without dataset-specific parameter updates. Pretrained on synthetic datasets drawn from a structural causal model prior, DiffPFN conditions its denoising process on observed training examples to generate samples with numerical and categorical features. Its transformer architecture jointly denoises all features, enabling a single pretrained model to operate across datasets with varying numbers and types of features. Experiments on eight datasets show competitive data fidelity, with a – speedup over diffusion baselines in total training and sampling time for new datasets. When fine-tuned, the pretrained model achieves higher downstream performance than a model trained from scratch at every budget up to 1,000 steps.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.