acceptodds
Under review as a conference paper at ICLR 2027

Gap-Aware Predictive Memory for Test Time Training

Abstract

Test-Time Training (TTT) expands recurrent sequence modeling by treating hidden states or KV-Binding as input-adaptive fast weights that continuously learn during inference. Existing TTT frameworks typically optimize these weights using self-supervised next-token prediction or full-state reconstruction. We argue that fitting fast memory directly to raw future representations is fundamentally redundant: a substantial portion of the target state is already explained by the frozen model’s pretrained predictive prior. Under bounded memory budgets, this redundancy forces fast weights to expend limited capacity on static, task-invariant dynamics rather than novel, context-specific interactions. To address this limitation, we introduce **GAP-TTT** (**G**ap-**A**ware **P**redictive Memory), a TTT framework that explicitly decomposes sequence prediction into a pretrained slow prior and an adaptive fast correction. Instead of modeling full target representations, **GAP-TTT** trains fast memory exclusively on the predictive gap, the semantic residual unaccounted for by the frozen backbone. Specifically, a lightweight semantic predictor exposes the slow model’s future latent expectations, and temporal consistency maps the resulting prediction error into both the memory write target and transition-aware key features. This unifies what to memorize and how to index it without requiring auxiliary inner-loop objectives. GAP-TTT updates these weights via online causal ridge regression, enabling hardware-efficient parallel prefix computation during inference, eliminating test-time backpropagation and iterative gradient descent. In the RULER and LongBench-v2 tests, GAP-TTT consistently outperforms other TTT baseline models while requiring only 25% of their training parameters, validating our argument that fast memory should focus on local contextual residuals.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.