acceptodds
Under review as a conference paper at ICLR 2027

Is your LLM a sequence model on the training history? The origins and consequences of anticipation

Abstract

Recent work showed that, in cyclic fine-tuning settings, Transformers exhibit striking anticipatory recovery: the loss on an upcoming document decreases before that document is revisited, even though models hold no explicit memory of training history. Despite early efforts to explain this behavior, our understanding of anticipation remains limited. We first show that anticipation is broader than previously demonstrated: beyond individual-documents, Transformers anticipate upcoming data domains, such as English and French, and the behavior extends to pre-training, even to models 5000 times smaller. Second, we show that anticipation can be interpreted as sequence model-like behavior. Anticipating the next domain leads to lower training loss, and this loss advantage realizes a substantial fraction of that achievable by sequence models that explicitly learn class order. To a limited extent, anticipation transfers even to second-order structures, where the next domain can only be predicted from the previous two. Third, we offer a potential explanation of anticipation as gradient alignment developing through an implicit bias of SGD in cyclically ordered training. The resulting picture helps account for the surprising optimizer dependence of anticipation: at matched held-out loss, AdamW develops strong anticipation, while Muon does not anticipate. Finally, we show that structured training order can affect what the model learns, and can act as a source of representational bias during optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.