acceptodds
Under review as a conference paper at ICLR 2027

Fast Weights in Fast Weights: Towards Bridging Parallel and Sequential Test-Time Learning

Abstract

Test-time training (TTT) layers store context in the weights of a small inner network that is trained by gradient descent as the sequence is read. For hardware efficiency, practical TTT evaluates every gradient in a chunk at the chunk-entrance weights, discarding the effect of earlier within-chunk updates on later gradients. What this mini-batch approximation loses, and which part is worth recovering, is not well understood. Using block-level backward error analysis, we show that the leading-order gap between sequential and mini-batch TTT is the causal interaction , which is also the leading perturbation of the inner loop's modified vector field. For a two-layer fast-weight MLP, splits into four layer-wise interactions with closed forms. This motivates , a fast-weights-in-fast-weights design that keeps the first layer as a chunk-level fast weight and makes the output layer a token-level memory, while retaining chunk parallelism. With features fixed, the output layer follows the delta rule, which a chunkwise WY solve computes exactly. The influence of earlier first-layer updates reduces to feature transport, a causal linear-attention pass that shifts the key and query features before the solve. For small step sizes, we prove that matches the sequential output-layer update through second order, leaving the two first-layer interactions as the entire second-order residual. From 340M to 1B parameters, improves on the chunk-parallel baseline LaCT in perplexity and average zero-shot accuracy at every scale. On long-context tasks, it improves the average score by 8.5 points over LaCT and by 3.8 points over an ablation without feature transport.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.