Selectivity Sets the Recall Range in Hybrid Language Models
Abstract
Hybrid language models interleave a few softmax-attention layers among many state-space model (SSM) layers. It has been unclear which properties of the SSM predict the query-key distances at which attention improves recall. Replacing every attention output with an input-independent constant isolates the SSM's autonomous recall. We show that attention sustains recall at distances where this recall has declined, a handoff that is clear in Granite-4.0, weak in Nemotron-H, and absent in one parallel- and one shared-attention hybrid. Input-dependent timestep gating (Δ-selectivity) sets the SSM's recall range (g50, the gap at which chance-corrected recall falls to half its peak). It does so by keeping interfering tokens from eroding the state. On a controlled two-key associative task, turning selectivity on extends the range 2.4-fold (from 24 to 58 tokens), whereas fixing a trained selective model's timestep at each head's mean drops peak recall below query-blind guessing. Turning selectivity on during training shifts the onset of an added attention layer's benefit to larger query-key gaps. Changes in recall range quantitatively predict these shifts. Across natively trained backbones, the range predicts onset on held-out architectures with a mean absolute error of 3.9 tokens, against 10.3 for the strongest architecture-plus-selectivity baseline. In Granite-4.0 and Nemotron-H, removing input-dependent gating significantly reduces the SSM's autonomous recall while attention-on recall stays well above it. Native SSM recall depends on this gating. On Granite, fixing each sequence's timestep at its native average shrinks the range from 11 to 3 tokens without a significant drop in peak recall. Timestep gating sets how far the SSM recalls on its own, and that range predicts the query-key distance at which attention begins to help.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.