acceptodds
Under review as a conference paper at ICLR 2027

RoleLex: Recovering Cross-Speaker Referential Structure from Frozen Language-Model States

Abstract

When two people talk about the same thing, the words they use for it drift: in one MapTask dialogue a route giver's "the old mill" is answered with "the mill wheel," and both speakers then settle on "it." A system that tracks such references has to store each mention as it arrives, before it knows what will be asked later. We study how to build that stored key from a frozen dialogue decoder and find that no single layer holds the evidence. On MapTask dialogues, decoder depths read out separately pick different antecedents for 47–68% of queries, and the best depth changes from one backbone to the next. RoleLex keeps the depths apart instead of choosing between them: four depths are projected and normalized independently and each occupies its own block of a 256-dimensional key, so one dot product between keys equals the average of four depth-wise cosines. On 13,077 human-linked mentions from 128 MapTask dialogues, evaluated leave-one-map-out, this fixed key beats the best development-selected single depth for Qwen3.5-9B, Gemma 4 E4B-it, and OLMo 3 7B, matches or beats the best selected pair of depths, and costs the same parameters and arithmetic as a final-layer key. Controls show that what transfers across backbones is keeping the depth scores separate; a dense map over the same states has four times the parameters, does worse, and uses fewer than half of its output directions. The gains are largest when no antecedent shares the query's wording, carry over to deciding whether any antecedent exists, and reappear on OneCommon.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.