acceptodds
Under review as a conference paper at ICLR 2027

A PAUSE IN TIME SAVES NINE: PUNCTUATION-GUIDED ATTENTION FOR EFFICIENT MULTI-TOKEN REASONING

Abstract

Standard self-attention scores each query against individual keys, which makes it difficult for Transformers to condition on the joint presence of multiple non-contiguous concepts. Prior remedies rely on fixed-size local neighborhoods that are not aligned with the semantic structure of text and scale poorly to long contexts. We introduce Punctuation-Guided Attention (PGA), which inserts learned summary tokens at punctuation boundaries (PAST) and restricts attention beyond a local sliding window to these summaries. This hierarchical pattern gives every token access to compressed representations of all preceding segments while filtering token-level noise. On a synthetic multi-token retrieval task, PGA maintains above 96% accuracy as block count and query length grow, where standard and summary-token-only Transformers collapse to near-zero accuracy. On real data, PGA improves convergence when pretraining from scratch and improves downstream accuracy under continued pretraining of Llama 3.2 1B. Because past standard tokens can be evicted once summarized, PGA also substantially reduces the inference-time key-value (KV) cache

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.