EchoStripe Attention: Cross-Layer Stripe Memory for Efficient Long-Context Inference
Abstract
Long-context inference is constrained by the quadratic cost of self-attention, whereas aggressive sparsification can remove structured long-range interactions needed by downstream tasks. We focus on lag-indexed diagonal stripes in the attention matrix: query–key pairs with the same relative offset share a common lag and can be represented compactly by a one-dimensional profile. Existing dynamic routing typically treats each layer as an independent decision, making the resulting sparse route sensitive to local fluctuations. We introduce EchoStripe, a training-free framework for stateful cross-layer sparse routing. Its Stripe Memory recursively propagates a compact preceding-layer lag-profile state and fuses it with current-layer evidence using a head-specific reliability-aware weight. EchoStripe further applies Mask Refinement, a bounded token-guided procedure that corrects the complete causal base mask through conservative additions and pruning. Across LongBench with Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct, and InfiniteBench with Llama-3.1-8B-Instruct, EchoStripe remains within 0.42, 0.85, and 1.16 points of the strongest sparse baseline, respectively. At 128K tokens, it reduces measured attention-operator latency by 5.20× relative to MInference and achieves a 7.73× speedup over Full Attention. On the common 20-dataset LongBench suite, Stripe Memory and Mask Refinement jointly improve the current-layer-only variant by 5.90 points. These results show that cross-layer routing state provides a favorable accuracy–efficiency operating point for long-context inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.