acceptodds
Under review as a conference paper at ICLR 2027

DisLA: Linear Attention with Distinct Use and Update Decisions

Abstract

Recurrent linear attention stores sequence history in a fixed-size state. Existing gated Delta models use the current input to correct key–value associations in the state, then read the updated state to form an output. Thus, the same correction is committed to the persistent state and contributes to the current output. However, persistent state updates need to account for future outputs, whereas the current output serves only the present step: the two roles may require different state corrections. We introduce DisLA, which learns separate corrections for state updates and current outputs. The two control projections share a base, with a low-rank branch representing their difference. We instantiate this mechanism in Gated DeltaNet-2 (GDN2), preserving its feed-forward width and recurrent state size while supporting chunkwise parallel training. Joint erase–write decoupling in DisLA-GDN2 adds only 0.17% parameters. In matched experiments with approximately 133M parameters and 3B valid next-token prediction targets from FineWeb-Edu per model, DisLA-GDN2 reduces WikiText-2 perplexity by 2.6% and improves average accuracy across LAMBADA and eight reasoning tasks by 1.22 percentage points over GDN2. It also raises the mean score across all 13 RULER tasks at 4K from 1.81 to 2.65.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.