Scalable Operator Learning via Tensorized Transformer
Abstract
Attention-based models are widely used to learn solution operators for partial differential equations (PDEs), but the computational cost of attention grows quadratically with the number of tokens. Windowed attention reduces this cost but limits direct interactions to local neighborhoods. Factorized attention attends along subsets of coordinates, but its approximation properties and practical benefits remain unclear. We study a tensorized Transformer that stacks independently parameterized blocks, each attending along a single coordinate, so that one sweep over all coordinates yields global interaction. Viewing attention as an integral operator, we approximate kernels by finite sums of tensor products of one-dimensional kernels. We then prove that tensorized Transformers approximate continuous solution operators uniformly on compact sets. We compare against full and windowed attention under matched model and training configurations on four two- and three-dimensional benchmarks from The Well. Across three model sizes, tensorized attention reduces VRMSE by 17.5% relative to these alternatives and runs 2.1 faster than full attention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.