LYNX: Linearizing Transformers with Timescale-Separated Memory
Abstract
Linearizing pretrained Transformers reduces long-context inference costs without pretraining from scratch, but preserving their capabilities under limited adaptation remains challenging. Our spectral analysis reveals low effective rank in the linearized baseline's recurrent state, with energy concentrated in a few directions. This limited use of the state's dimensions motivates a memory design with distinct retention timescales. We propose LYNX, a lightweight linearization method that combines sliding-window attention for exact recent access with timescale-separated recurrent states for historical information. The states jointly support association updates and fact retention, while a controller decouples retention from write allocation. A Landmark cache complements recurrent memory with exact access to selected historical tokens. Our analysis characterizes how retention shapes spectral concentration and gives sufficient conditions for the separated states to attain higher effective rank than a shared-retention baseline. Experiments on 8B-parameter LLaMA-family models evaluate general and long-context capabilities. LYNX achieves average accuracy across six reasoning and knowledge tasks, demonstrating general capability recovery. On long-context tasks, LYNX reaches a seven-task LongBench average of 26.07 and outperforms the compared linearization baselines across all eight Extended RULER task groups at 4K.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.