acceptodds
Under review as a conference paper at ICLR 2027

Approximation and Generalization Theory of Neural Operator Transformers for Fr\'echet-H\"older Operator Class

Abstract

While Transformer-based networks have exhibited remarkable performance in operator learning, the theoretical analysis remains underdeveloped. In this paper, we establish a comprehensive approximation and generalization theory of neural operator Transformers for learning operators in the Fr 'echet-H "older class. This class extends the Lipschitz continuous and times Fr 'echet differentiable classes commonly considered in operator learning. For target operators in the considered Fr 'echet-H "older class, we derive uniform approximation bounds on Sobolev balls of arbitrary radius, explicitly specifying network size required for a prescribed accuracy. Under a Sobolev sub-exponential assumption on the input distribution, we establish non-asymptotic generalization bounds for the neural operator Transformer estimator. By proving a matching lower bound, we demonstrate that neural operator Transformers achieve the minimax optimality over the Fr 'echet-H "older class. We further extend our analysis to practical settings where input and output functions are observed only on random grids, providing explicit error bounds in the number of training functions and the number of per-function observations. Numerical experiments on PDE solution operators further verify and support our theoretical analysis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.