Test-Time Training for Pretrained LLMs: An Empirical Study of Limited Fast-Weight Utilization
Abstract
Test-time training (TTT) enables a model to learn from its input context during inference by updating a subset of its parameters, effectively turning them into a form of dynamic memory. In-Place TTT extends this approach to pretrained LLMs by repurposing existing MLP parameters as fast weights through continued pretraining. Unlike dedicated memory trained from scratch, these parameters already encode useful computation, creating a tension between contextual adaptation and pretrained capabilities. We study this tension through a detailed analysis across multiple fast-weight learning rules, varying the update objective as well as write and forget gating mechanisms. Our central finding is that performance gains can arise largely from changes learned by the underlying model during TTT-aware training, while the direct contribution of inference-time fast-weight updates is often surprisingly small. To understand how this behavior emerges, we examine training dynamics and find that TTT-aware training progressively limits how strongly fast-weight updates alter the model’s computation, through mechanisms that depend on the learning rule. Amplifying updates at inference time can leave long-context retrieval intact while degrading pretrained capabilities, suggesting that capability preservation may partly explain this suppression. We further analyze how TTT-aware training shapes context utilization: the underlying model learns to attend more strongly to relevant evidence even without fast-weight updates, whereas the write gates governing those updates become more selective and focused on query-relevant tokens despite their limited direct contribution. Together, our results highlight the challenges of turning pretrained parameters into dynamic memory without disrupting the computation they already encode.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.