Modular TTT: Rethinking Test-Time Training as Composable Modules
Abstract
Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically hard-code each variant separately, which makes it difficult to design new TTT methods and to isolate the role of each component. To address this, we propose Modular TTT, a framework that represents the inner learner as a directed acyclic graph and exposes the fast-weight network, loss function, learning rate, weight decay, and normalization as explicit design dimensions. Once the forward pass of each fast-weight primitive is specified, Modular TTT automatically composes the corresponding backward and query-view computations over the graph. Using Modular TTT, we systematically ablate the components of TTT and find that the learning rate, weight decay, a single-layer nonlinearity, and a dot-product loss improve performance. By contrast, deeper fast-weight networks and normalization tend to hurt performance because they induce excessively large activations, while residual connections, gating, and MSE loss provide little measurable benefit. Guided by these findings, we train the best resulting variant as 410M- and 1.45B-parameter models on 100B tokens, and observe training loss and benchmark performance comparable to Gated DeltaNet.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.