DisLA: Linear Attention with Distinct Use and Update Decisions
Abstract
Recurrent linear attention stores sequence history in a fixed-size state. Existing gated Delta models use the current input to correct key–value associations in the state, then read the updated state to form an output. Thus, the same correction is committed to the persistent state and contributes to the current output. However, persistent state updates need to account for future outputs, whereas the current output serves only the present step: the two roles may require different state corrections. We introduce DisLA, which learns separate corrections for state updates and current outputs. The two control projections share a base, with a low-rank branch representing their difference. We instantiate this mechanism in Gated DeltaNet-2 (GDN2), preserving its feed-forward width and recurrent state size while supporting chunkwise parallel training. Joint erase–write decoupling in DisLA-GDN2 adds only 0.17% parameters. In matched experiments with approximately 133M parameters and 3B valid next-token prediction targets from FineWeb-Edu per model, DisLA-GDN2 reduces WikiText-2 perplexity by 2.6% and improves average accuracy across LAMBADA and eight reasoning tasks by 1.22 percentage points over GDN2. It also raises the mean score across all 13 RULER tasks at 4K from 1.81 to 2.65.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.