acceptodds
Under review as a conference paper at ICLR 2027

Mechanisms Off the Data Manifold: Distribution-Relative Identifiability in Neural Operators

Abstract

The standard Navier-Stokes benchmark for operator learning contains a blind spot introduced by its own data pipeline: the generator enforces zero spatial-mean vorticity across all examples, whereas the true solution operator acts as the identity along this constant mode. Operators trained to benchmark accuracy recover that response at roughly one-eighth of its true strength, incurring a 28% error under a mean offset that the true dynamics handle trivially. This failure is invisible to standard in-distribution evaluations, is predictable from the training covariance alone, and admits an exact, zero-cost structural repair once diagnosed. This reflects a broader phenomenon: function-value empirical training identifies local mechanisms only along directions excited by data leaving the remaining modes governed entirely by architectural inductive bias. For translation-equivariant model families, this completion is predictable from a continuity dichotomy between localized spatial kernels and per-mode spectral weights. Controlled experiments show that failure follows excitation rather than frequency, that common canonicalization removes the evidence required to resolve a symmetry's mechanism, and this vulnerability amplifies as physical dissipation weakens. The diagnosis is actionable: training covariance flags which data to acquire, while a known symmetry enables exact structural repairs that degrades gracefully under misspecification.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.