Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression
Abstract
Mechanistic studies of in-context learning (ICL) have connected transformer computations to iterative optimization algorithms for linear regression and related tasks, often using linear or ReLU attention variants. For nonlinear ICL, prior work has related softmax and kernelized attention to functional gradient descent, but it remains unclear whether a standard transformer with softmax attention can implement a convergent solver with an end-to-end prediction-error guarantee. In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass. Under bounded-data assumptions, we construct a single-head transformer whose forward pass approximately implements preconditioned Richardson iteration on the associated kernel system. The construction uses blocks and MLP width to achieve -accurate prediction for prompts of length . Our construction reveals a functional decomposition within the transformer architecture: softmax attention produces a row-normalized Gaussian-kernel operator needed for cross-token interactions, while MLP layers act locally to approximate the intra-token scalar arithmetic required by the update. Empirically, we train GPT-2-style transformers on Gaussian-process regression tasks and observe that they progressively align with the exact Gaussian KRR estimator across depth in terms of both prediction error and induced weights, with ablations further supporting this trend. Comparisons with classical KRR solvers also show that deeper layers align with later solver iterates. Together, we empirically demonstrate that the pretrained transformers exhibit progressive refinement toward exact Gaussian KRR across depth, and theoretically establish inexact preconditioned Richardson iteration as a concrete mechanism for approximating this predictor within an explicitly constructed softmax-attention transformer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.