acceptodds
Under review as a conference paper at ICLR 2027

How Do Language Models Route Information Through Repeated Tokens?

Abstract

Repeated words are common in natural-language inputs to Language Models (LMs). We assume they serve as the anchors where the contextual information is collected and stored, and thereby detected and forwarded to the later repetition, by a set of attention heads named Duplicate Token Heads (DTHs). Therefore, we investigate such attention heads and find that: (1) Not all DTHs necessarily route information that is causally relevant to a given task since only a subset of attention heads can access the appropriate task subspace and route that information. (2) DTHs do not identify information-routing targets through simple input-token matching alone; contextual information also plays an important role. (3) DTHs become increasingly important as context length grows: ablating them causes substantially greater accuracy degradation in long contexts than in short ones, suggesting that direct local processing can serve as a competitive fallback only for short contexts, whereas long-context inference increasingly relies on anchor-based information. Our findings provide preliminary evidence that LMs utilize a cache-and-forward mechanism for information from earlier context, rather than re-traversing preceding tokens at later positions to retrieve that information.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.