When Are Physics Gradients Corrective? Update-Space Dependence in Autoregressive Neural PDE Operators
Abstract
Autoregressive neural PDE predictors can remain locally accurate under a physical-coefficient shift while their long rollouts deteriorate. Governing-equation residuals offer label-free test-time objectives, but a small residual does not establish that its gradient will correct predictive error, nor does a corrective direction guarantee that an optimizer can use it. We separate these questions in controlled viscous Burgers and forced Navier–Stokes systems. Our central experiment holds real update-space dimension fixed. In two Fourier-operator training regimes, architecture-defined pointwise ((d=648)) and output-head ((d=1,249)) spaces have larger mean horizon-100 corrections than ten frozen random spaces at the same size. A two-way model-seed/initial-condition bootstrap supports the pointwise effect under both one- and five-step protocols; the head effects remain positive in mean but statistically inconclusive. A structured (d=4,096) counterexample rules out a blanket preference for structure. Complementary interventions separate scalar residual value from gradient direction: pooled residual–error association is near zero but heterogeneous across seeds, while equal-norm physics updates rank above matched random motion and reversal in aggregate. Rank-four adapters can nevertheless have high alignment and almost no rollout gain. A well-fit non-spectral DilResNet also exhibits positive alignment, directional specificity, and a 12.10% mean physics-only adaptation gain. Correctability is therefore a joint property of the objective and accessible update space, not residual magnitude, cosine, or dimension alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.