acceptodds
Under review as a conference paper at ICLR 2027

Synthetic to SPX: Model-Based Pre-Training for Diffusion Models of Financial Time Series

Abstract

Diffusion models have achieved remarkable success in image, video, and text generation. We apply them to financial time series. The main obstacle is the scarcity of data: only one realized price trajectory is available for each asset. Finance has, however, a major advantage: fifty years of mathematical models that capture the main mechanisms of markets and can be simulated at will. These models provide unlimited training data, while their known statistical properties enable controlled evaluation of generated trajectories. We therefore propose model-based pre-training: first, train the score network on parametric models of increasing complexity; second, fine-tune it on the single real trajectory. In the first phase, we compare two approaches: one learns the joint path distribution in , the other factorizes it into conditional distributions over . The autoregressive formulation is the most data-efficient, reaching a given fidelity with up to four times fewer trajectories; with abundant data, all architectures perform alike. We then fine-tune on the overlapping windows of a single trajectory, first on a path-dependent volatility model calibrated on SPX and held out from pre-training, where the ground truth is known, and then on SPX itself. Pre-training improves adaptation in this low-data regime, and what matters is the proximity of the source models to the target rather than their diversity. On SPX, the pre-trained model recovers volatility clustering, heavy tails and the leverage effect, which training from scratch misses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.