acceptodds
Under review as a conference paper at ICLR 2027

LYNX: Linearizing Transformers with Timescale-Separated Memory

Abstract

Linearizing pretrained Transformers reduces long-context inference costs without pretraining from scratch, but preserving their capabilities under limited adaptation remains challenging. Our spectral analysis reveals low effective rank in the linearized baseline's recurrent state, with energy concentrated in a few directions. This limited use of the state's dimensions motivates a memory design with distinct retention timescales. We propose LYNX, a lightweight linearization method that combines sliding-window attention for exact recent access with timescale-separated recurrent states for historical information. The states jointly support association updates and fact retention, while a controller decouples retention from write allocation. A Landmark cache complements recurrent memory with exact access to selected historical tokens. Our analysis characterizes how retention shapes spectral concentration and gives sufficient conditions for the separated states to attain higher effective rank than a shared-retention baseline. Experiments on 8B-parameter LLaMA-family models evaluate general and long-context capabilities. LYNX achieves average accuracy across six reasoning and knowledge tasks, demonstrating general capability recovery. On long-context tasks, LYNX reaches a seven-task LongBench average of 26.07 and outperforms the compared linearization baselines across all eight Extended RULER task groups at 4K.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.