acceptodds
Under review as a conference paper at ICLR 2027

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent-Depth Language Models

Abstract

Recurrent-depth language models repeatedly apply shared network blocks to refine latent representations, enabling additional test-time computation without generating explicit intermediate reasoning tokens. However, each recurrent step recomputes full self-attention over the entire context, repeating global attention routing whose cost is quadratic in context length. We study how attention routing evolves across recurrent depth and uncover a consistent separation in convergence timescales: routing-related quantities, including the attention support and the attention distribution, stabilize substantially earlier than representation-related quantities such as hidden states and attention outputs. This suggests a two-stage structure in recurrent inference, where the model first discovers a sparse working set of relevant context and then continues refining representations over largely the same routing support. Motivated by this structure, we introduce WISE (Working-set Inference with Support Exploitation), a training-free method that uses unrestricted global attention during early recurrent steps to discover a block-structured working set, and then reuses it as the routing support in later steps while keeping recurrent refinement and within-support attention computation dynamic. Controlled interventions show that discovering the working set over multiple recurrent steps, rather than from the first step alone, yields more effective working sets, and that reusing only the routing support better preserves model behavior than more restrictive forms of late attention reuse. Across multi-hop QA benchmarks, WISE largely preserves the behavior of full-attention inference, while matched context-scaling experiments reveal an increasingly favorable quality–efficiency tradeoff as the routing support becomes increasingly sparse with longer contexts. A sparse-attention implementation translates this structured sparsity into practical GPU acceleration, achieving up to a late-step attention speedup over native FlashAttention at 4K context.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.