acceptodds
Under review as a conference paper at ICLR 2027

Language models use both proactive and retroactive mechanisms to make compositional inferences

Abstract

One of the most powerful features of modern language models (LMs) is their ability to generalize about novel in-context information. A fundamental component of this ability is compositional inference: the capacity to combine multiple new relations such as 'Tom's father is Steve' and 'Jim's father is Tom' in order to arrive at an unstated inference such as 'Jim's grandfather is Steve'. In this work, we study how LMs make these kinds of inferences. Using a variety of causal interventions on Llama 3.1 405B Instruct, we first demonstrate that 405B Instruct uses two functionally independent mechanisms to do so: a proactive mechanism that makes generalizations as soon as information is available in-context, and a retroactive mechanism that reasons about what is in-context only when it is necessary for predicting the immediate next token. We then consider why 405B Instruct uses two different mechanisms to make compositional inferences. We find that 405B Instruct uses contextual information to choose when to use the proactive mechanism, and that the proactive mechanism enables it to make deeper inferences than what the retroactive mechanism supports on its own.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.