Differential Attention for State Space Models
Abstract
The Transformer architecture has emerged as the dominant paradigm for sequence modeling; however, it suffers from quadratic complexity in sequence length. In parallel, state space models (SSMs), such as Mamba, have been proposed as efficient alternatives, offering sub-quadratic complexity while maintaining competitive performance. The Differential Transformer demonstrated that transformers tend toward attention over-allocation, which can impair retrieval and long-context reasoning. They address over-allocation by subtracting attention maps, thereby improving signal selectivity and mitigating noise. Existing attempts to transfer this idea to SSMs rely on parallel reduced Mamba blocks that incur limited scalability, and significant computational overhead. We introduce a generalized formulation of differential computation for state space models derived directly from structured state space duality (SSD). Rather than constructing parallel SSM branches, our formulation shows that differential computation naturally emerges as an operation on state-space readout. This perspective enables both an exact single-pass formulation and an efficient single differential circuit implementation without modifying the underlying SSM kernels. We instantiate the proposed framework in both Mamba-2 and Mamba-3, yielding Diff-Mamba-2 and Diff-Mamba-3. Across language modeling benchmarks, our models consistently outperform their corresponding baselines by up to 1.5 average points, while preserving the computational efficiency of modern SSM architectures. We further demonstrate improved retrieval on real-world long-context benchmarks, and sharper token-level interaction patterns through interpretability analysis. These results establish differential computation as a principled and effective design paradigm for state space models. Our code is publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.