acceptodds
Under review as a conference paper at ICLR 2027

Scalable Operator Learning via Tensorized Transformer

Abstract

Attention-based models are widely used to learn solution operators for partial differential equations (PDEs), but the computational cost of attention grows quadratically with the number of tokens. Windowed attention reduces this cost but limits direct interactions to local neighborhoods. Factorized attention attends along subsets of coordinates, but its approximation properties and practical benefits remain unclear. We study a tensorized Transformer that stacks independently parameterized blocks, each attending along a single coordinate, so that one sweep over all coordinates yields global interaction. Viewing attention as an integral operator, we approximate kernels by finite sums of tensor products of one-dimensional kernels. We then prove that tensorized Transformers approximate continuous solution operators uniformly on compact sets. We compare against full and windowed attention under matched model and training configurations on four two- and three-dimensional benchmarks from The Well. Across three model sizes, tensorized attention reduces VRMSE by 17.5% relative to these alternatives and runs 2.1 faster than full attention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.