acceptodds
Under review as a conference paper at ICLR 2027

One Attention Head Makes Where to Look Almost Free: The Cost of Unknown Support with and without Adaptive Selection

Abstract

A transformer must decide where to look before it decides what to compute. We measure the statistical price of that decision in nonparametric regression. The target depends on an unknown set of of its input coordinates, one token each, through a smooth function of variables, namely an anisotropic Besov ball with a sup-norm envelope; the design is i.i.d. with a density bounded above and below uniformly in and , the noise is Gaussian, and is fixed. Without adaptive selection the price is combinatorial. Every estimator that is affine in the training responses, with offset and coefficients fixed before the responses are read, as in kernel ridge regression at a label-independent regularisation level, random-feature ridge, or lazy/NTK rules, has worst-case risk at least with , which does not vanish once ; for integrability index the exponent is the smaller, shifted one. One positional attention head makes the price additive and polylogarithmic in . The exact clipped least-squares minimiser over an explicitly budgeted class of one-head transformers with positional keys attains in the effective smoothness , up to logarithmic factors in and with a constant free of , plus a single additive selection term polylogarithmic in . The mechanism is a whitening of the positional codes: for every support, the head can attend equally to the relevant tokens, with only a controlled remainder elsewhere, inside one parameter budget that does not depend on which tokens they are. At with all structural parameters fixed, the response-affine worst-case risk stays bounded away from zero while the attention bound vanishes. The saving is not specific to attention, since a direct search over supports attains the same leading bound; what attention supplies is a compact, continuous realisation of it. The upper bound concerns the exact minimiser, and no optimisation guarantee is claimed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.