acceptodds
Under review as a conference paper at ICLR 2027

ASIR: An Agent-Structure-Aware Indexer with Block-Local Reuse for Native Sparse Attention

Abstract

Long-context agents accumulate conversations, tool outputs, and files, increasing the cost of attention over the growing history. Native sparse attention uses an Indexer to select a fixed number of tokens for attention, but selection still scans the full context. Hierarchical Indexers narrow the search by selecting a prescribed number of blocks before scoring their tokens. However, their budgets overlook request-specific differences in evidence coverage, and they rescan every selected block at each decoding step. We present ASIR, an Agent-Structure-Aware Indexer with block-local candidate reuse. During prefill, ASIR combines native segment boundaries, request relevance, and spatial extent to allocate candidate capacity. During decoding, it retains historical candidates in some selected blocks and refreshes others, scoring all candidates with the current query. The two components reduce the spatial extent and temporal repetition of Indexer scoring, respectively. On LongBench and LongMemEval, ASIR achieves aggregate quality comparable to a hierarchical baseline searching one quarter of the context while reducing both candidate-search capacity and repeated token-level scoring. Across 16K–128K contexts, ASIR achieves – Indexer speedup over native sparse attention and – over this hierarchical baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.