Virtual-Transition Temporal-Difference Learning for Autoregressive PDE Dynamics Models
Abstract
Autoregressive PDE models are typically trained to predict each observed next state from its input window. We study a complementary training choice: which examples should exchange prediction-error information? Building on the Markov reward process formulation for supervised learning, we use virtual-transition temporal-difference (TD) learning to couple residuals. These transitions need not follow the physical dynamics; our default pairs different trajectories at the same time index and leaves the prediction architecture and inference procedure unchanged. For linear models, we quantify the cost of imperfect inverse-covariance weighting: the bound on excess covariance relative to oracle generalized least squares (GLS) is quadratic near exact matching, and the same bound controls leading-order finite-horizon rollout mean squared error (MSE). Strict leading-order improvements over supervised least squares persist for exact rollouts at sufficiently small target noise. A local nonlinear extension characterizes when TD retains GLS efficiency and how weighting mismatch affects rollout error. Controlled experiments verify these theoretical results. Using the same pairing rule across PDE domains, TD improves one-step test accuracy on domains and short-horizon rollout accuracy on , with median error reductions of and ; the median reduction over steps 21–60 is across the eligible domains. Original-scale experiments and pairing, backbone, and multistep studies establish the method's utility across several model and training configurations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.