Training-Free Universal Approximation by Prompting Random Transformers
Abstract
How expressive is prompting a transformer? Answering this question is important for disentangling the roles of prompting, architecture, and pretraining in transformer models, and for determining whether task-specific behavior must be stored in model weights or can instead be induced at inference time through the prompt. We show, in an approximation-theoretic sense, that pretraining is not strictly necessary: a single-layer softmax attention network with random, untrained weights can approximate any Hölder function on a compact manifold when steered by an appropriate soft prompt. Building on the connection between softmax attention and kernel methods, we construct, for each target function, an explicit query-independent soft prompt by solving linear systems that match attention logits to Gaussian kernel exponents, enabling the frozen transformer to emulate the classical Nadaraya–Watson kernel estimator. This construction requires only a mild rank condition on the weights, which we show holds almost surely under Gaussian initialization. The prompted network inherits the theoretical guarantees of kernel regression, leading to universal approximation results with minimax-optimal rates governed by the manifold’s intrinsic dimension. Moreover, we quantify the cost of prompting, revealing a tradeoff among the norms of the constructed prompt tokens, prompt length, and hidden dimension. Numerical experiments corroborate the constructions and predicted rates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.