acceptodds
Under review as a conference paper at ICLR 2027

Integral DeltaNet: Improved Linear Attention with Adaptive Memory Integration

Abstract

Linear attention replaces softmax attention with a linear formulation, enabling linear-time computation with a fixed-size recurrent state instead of a growing key–value cache. These computational and memory advantages have attracted growing interest in linear attention for long-context language modeling. However, compressing past information into a fixed-size state causes *memory collisions* that hinder long-context understanding. Although DeltaNet mitigates memory collisions through delta-rule updates and Gated DeltaNet extends this approach with gating, these mechanisms may also discard information relevant to future predictions. To address this limitation, we introduce *Integral DeltaNet*, which maintains two complementary recurrent states: a fast-updating *delta state* that captures current key–value associations and an *integral state* that preserves a weighted history of the delta state. At readout, the model adaptively combines these states to access information across different timescales. We further derive a chunkwise parallel formulation that enables hardware-efficient training. We evaluate Integral DeltaNet using 340M and 1.3B parameter models trained from scratch on 15B and 100B FineWeb-Edu tokens, respectively. Comparisons with strong recurrent baselines show overall improvements in language modeling, commonsense reasoning, real-world and synthetic retrieval, long-context understanding, and length extrapolation. Training throughput remains close to that of the backbones.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.