acceptodds
Under review as a conference paper at ICLR 2027

In-Context Regret: Sequence Models Are Online Learners

Abstract

The search for post-transformer architectures has converged on test-time-learning layers—linear attention, DeltaNet, gated variants, Titans, mesa/ridge layers—each from a hand-chosen objective and update rule. What is missing is a performance calculus: a way to say, before training at scale, how good a memory rule is and how it must fail. We supply one. Every causal sequence-mixing layer is an online learner over its own key–value stream, and its in-context regret organizes the design space: linear attention is Hebbian regularized follow-the-leader ( under bounded predictions), DeltaNet is online gradient descent (), gating is discounting, and the mesa layer is exactly the Vovk–Azoury–Warmuth forecaster (). Against the memorizing comparator that attention estimates, we prove a memory–regret tradeoff: any layer whose state passes an -bit bottleneck suffers per-query recall error at least , with the binary entropy—while its regret against the linear comparator can be zero. What an architecture pays in state is the size of the comparator class it can chase. The theory is generative: parameter-free and strongly-adaptive online learning yield two layers without learned step-size or time-scale gates, and regret probes measure trained networks label-free. On enwik8 from 14M to 340M parameters it pays off twice. The coin-betting layer’s gap to the tuned delta rule closes monotonically with scale and flips sign: at 340M it posts the best test bpc of any recurrent rule, at 2–3× the delta and gated baselines' throughput. And length extrapolation mirrors the hierarchy at every scale—bounded-regret rules evaluate below their training-length bpc at 16× context, while Hebbian’s interference diverges.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.