acceptodds
Under review as a conference paper at ICLR 2027

Modular TTT: Rethinking Test-Time Training as Composable Modules

Abstract

Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically hard-code each variant separately, which makes it difficult to design new TTT methods and to isolate the role of each component. To address this, we propose Modular TTT, a framework that represents the inner learner as a directed acyclic graph and exposes the fast-weight network, loss function, learning rate, weight decay, and normalization as explicit design dimensions. Once the forward pass of each fast-weight primitive is specified, Modular TTT automatically composes the corresponding backward and query-view computations over the graph. Using Modular TTT, we systematically ablate the components of TTT and find that the learning rate, weight decay, a single-layer nonlinearity, and a dot-product loss improve performance. By contrast, deeper fast-weight networks and normalization tend to hurt performance because they induce excessively large activations, while residual connections, gating, and MSE loss provide little measurable benefit. Guided by these findings, we train the best resulting variant as 410M- and 1.45B-parameter models on 100B tokens, and observe training loss and benchmark performance comparable to Gated DeltaNet.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.