Approximation Theory for Transformers
Abstract
We establish the first comprehensive approximation theory for positional transformers that parallels classical neural network approximation theory. (i) We prove universal approximation uniformly on compact sets; in \(L_\mu^p\); uniformly on compact sets together with derivatives up to order \(k\); in weighted Sobolev norms. (ii) We establish , Sobolev, , and Barron approximation rates. (iii) The rate is optimal under continuous parameter selection. All these results follow from an exact realization theorem: we show that a positional transformer acting on tokens realizes suitable families of neural networks. This allows us to transfer classical neural network approximation theory, preserving approximation-rate exponents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.