acceptodds
Under review as a conference paper at ICLR 2027

Where Does the Recurrence Help? Retrieval-Distance Attribution in Hybrid Sequence Models

Abstract

Attention costs quadratic time in sequence length, and the two standard remedies fail differently. A sliding window bounds cost but cannot see past its edge — a positional limit. A compressed recurrence bounds cost but retains only a lossy summary of the prefix — a capacity limit. Hybrid architectures combine both, but evaluate with aggregate quality metrics that average the two regimes into a single number revealing neither. We present HALA, a block that computes sliding-window attention and a low-rank gated recurrence in parallel, mixes them with a learned per-head gate, and carries no position-wise feed-forward sublayer. Our contribution is less the architecture than the measurement. We stratify tokens by how far back their antecedent lies and attribute each path's contribution directly. Against an ablation with the recurrence removed, HALA is statistically indistinguishable on antecedents inside the 64-token window and consistently better outside it, by , and perplexity across three domains — and this advantage grows monotonically with scale, from at M non-embedding parameters to at M. On associative recall the separation from pure recurrence is categorical: Mamba-2 sits at chance across a full learning-rate sweep where HALA reaches . We are equally explicit about the limits. HALA does not lead aggregate perplexity in any domain: Transformer and Jamba are ahead throughout. Against Mamba-2 the comparison is a crossover with a legible mechanism — HALA leads on antecedents inside the window, Mamba-2 on those beyond it, and Mamba-2 carries the recurrent state per layer. Length-extrapolation stability is inherited from recurrent state rather than contributed, and all results sit at –M non-embedding parameters.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.