acceptodds
Under review as a conference paper at ICLR 2027

The Role of Attention Layers and Context Tokens in Entity Copying

Abstract

Large language models (LLMs) perform reliably on the task of entity copying, in which the model copies the tokens referring to an entity, or entity tokens, from the prompt into its outputs to answer a question. Although entity copying is straightforward for most LLMs, existing research does not provide a systematic account of which layers specialize in this fundamental task or how other tokens in the same sequence, termed context tokens, influence the model's ability to copy the entity tokens. To address these questions, we conduct experiments on Qwen3-8B using two novel methodologies: genie-in-a-bottle, which controls exactly which layers can participate in an entity copying task, and attention lobotomy, which cuts off specific tokens' attention to entity tokens without affecting the remaining attention distribution. We find that layers in the second half of the model are both necessary and sufficient for entity copying, whereas those in the first half play little to no role. Moreover, context tokens' joint attention to entity tokens plays a decisive role in successful entity copying, even though these context tokens do not store entity information in their intermediate states unless they satisfy particular semantic properties. Our findings establish the critical role of late layers in entity copying under the decisive guidance of context tokens, calling for future work on how entity information is propagated and consumed in models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.