acceptodds
Under review as a conference paper at ICLR 2027

Approximation Theory for Transformers

Abstract

We establish the first comprehensive approximation theory for positional transformers that parallels classical neural network approximation theory. (i) We prove universal approximation uniformly on compact sets; in \(L_\mu^p\); uniformly on compact sets together with derivatives up to order \(k\); in weighted Sobolev norms. (ii) We establish , Sobolev, , and Barron approximation rates. (iii) The rate is optimal under continuous parameter selection. All these results follow from an exact realization theorem: we show that a positional transformer acting on tokens realizes suitable families of neural networks. This allows us to transfer classical neural network approximation theory, preserving approximation-rate exponents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.