HotPOT: Operator Transformer with Heterogeneous Physics Experts for Large-Scale PDE Pre-training
Abstract
Large-scale PDE pre-training trains a single model on the data from multiple partial differential equations (PDEs). However, adding heterogeneous PDEs with different governing equations, physical variables, and variable interactions does not always improve predictive performance across tasks. Existing PDE mixture-of-experts models mitigate cross-task interference through learned routing, but typically use homogeneous expert architectures, making it difficult to directly reflect differences in state evolution across equations. We introduce HotPOT, an operator transformer with heterogeneous physics experts for large-scale PDE pre-training. HotPOT uses a shared encoder to capture information common across PDEs, while each PDE family uses a neural–numerical expert that pairs a learnable predictor with a fixed numerical update containing no trainable parameters. Each predictor outputs the quantities required by its paired numerical update, which then computes the next state. This design allows different PDEs to share a common encoder while retaining their own learnable predictors and numerical updates. Experiments across five PDE families and six settings show that HotPOT outperforms MoE-POT in all 24 matched-scale comparisons before and after fine-tuning, reducing rollout error by 50.7% on average. Notably, the 17M model also surpasses substantially larger MoE-POT models, and the performance gains extend to zero-shot prediction under unseen physical conditions. Code is available at https://anonymous.4open.science/r/HotPOT-code-1727/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.