acceptodds
Under review as a conference paper at ICLR 2027

Understanding and Accelerating Long-Context Extrapolation in Hybrid Models via State-Equilibration Curriculum

Abstract

Long-context adaptation of hybrid state-space model (SSM)–Transformer language models remains poorly understood: while multi-stage context-length training is effective in practice, it is unclear how long a model should remain at each intermediate context length before advancing to a longer sequence length. We find that long-context adaptation exhibits a two-phase behavior. When first exposed to a longer context, recurrent state dynamics rapidly stabilize: distance-dependent state drift sharply decreases and then plateaus, which we call state equilibration. Continued training beyond equilibration can still improve performance at the current context length, but these gains increasingly specialize to that length and transfer poorly to longer sequences. Based on this observation, we introduce State-Equilibration Curriculum (SEC), which advances to the next context length once recurrent-state dynamics stabilize. To support SEC efficiently, we develop a context-parallel implementation for hybrid SSM–Transformer models that removes serial cross-device state propagation while exposing chunk-boundary states as a low-overhead training signal. On Nemotron-H-8B and Bamba-9B-v2, SEC matches or improves fully trained multi-stage schedules while using about 30% fewer training tokens and reducing wall-clock training time by up to 42%. Our results suggest that intermediate context stages need not be trained to task convergence: advancing once state dynamics equilibrate provides a more efficient mechanism for multi-stage long-context adaptation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.