How Language Models Combine Multiple Binding Mechanisms
Abstract
Language models must bind entities together in context to later retrieve relevant information, for example when remembering which city a person lives in. Prior work has established three mechanisms (positional, lexical, and reflexive) that models use for this retrieval, but left open where the underlying binding signals are stored and how they are combined into an answer. In this work, we isolate the low-rank subspaces of the final token residual stream used by each mechanism and show that they can be individually manipulated. Further, we show that the subspaces contribute additively to the query key product of a single dominant attention head driving answer retrieval. Interventions on individual subspaces selectively redirect the head's attention and the model's final output. Finally, we examine the subspaces themselves: the positional subspace primarily encodes fact order, while the lexical and reflexive mechanisms share a semantic identity channel. Taken together, these findings demonstrate how known retrieval heuristics are represented and combined. More broadly, the work provides insight into how redundant mechanisms are implemented and reconciled within language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.