acceptodds
Under review as a conference paper at ICLR 2027

Test-time Memory From The Online Learning Perspective: Associative Prediction Game and Regret Optimality

Abstract

Test-time training (TTT) layers and Delta-style memories learn key-value (KV) associations as a sequence arrives. We study which update rules learn these association well and when the resulting memory can answer queries that differ from its update keys. We measure learning through regret, which is the extra prediction loss relative to a fixed comparison map. For exact recursive least squares (RLS), we connect prediction regret and retrieval error through the same regularized key covariance. Specifically, ordinary RLS has logarithmic regret when key norms and prediction residuals remain uniformly bounded. A classical Vovk–Azoury–Warmuth (VAW) readout achieves the optimal worst-case rate in the stated bounded prediction game without the residual condition. A separate construction shows that constant-step Delta can incur much larger regret. Our retrieval guarantees identify additional conditions on query coverage and how values and noise are generated. Furthermore, we analyze approximate updates and memories that forget old observations. Extensive experiments on WikiText streams show that RLS performs better in the evaluated stable settings, while resetting and discounting old information helps against controlled drifts. Meanwhile, a LongBench-v2 setting shows that selected Delta reconstructs native attention slightly better, while RLS performs better when attention is restricted to the stored context. These findings distinguish guarantees for learning KV associations from performance on query-time reconstruction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.