acceptodds
Under review as a conference paper at ICLR 2027

Pretraining Design Drives Gains in Zero-Shot PDE Rollouts

Abstract

Generalization beyond the pretraining distribution remains a major challenge for pretrained neural PDE solvers. We study zero-shot out-of-distribution (OOD) rollout prediction in fluid dynamics, where a pretrained model must predict future trajectories from a single initial condition without downstream finetuning, contextual demonstrations, or test-time adaptation. We uncover a surprisingly large performance gap: a lightweight ViT-style spatiotemporal model with only about 25M parameters outperforms the 629M-parameter pretrained Poseidon-L on several OOD systems, with approximately - lower error on two of them, despite its weaker in-distribution accuracy. We examine how pretraining design affects zero-shot transfer through controlled comparisons of autoregressive rollout, teacher forcing, non-autoregressive full-trajectory prediction, and different temporal conditioning schemes. We find that training formulation can change OOD errors by large factors. With full-history attention, teacher forcing improves our ViT's source-domain accuracy and yields better overall source-domain and OOD performance than AR rollout training, despite the training-inference mismatch. Comparisons with FNO reveal similar benefits in source-domain prediction and on several OOD systems, showing that the advantages of teacher forcing extend beyond our ViT architecture.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.