acceptodds
Under review as a conference paper at ICLR 2027

A Mechanistic Analysis of In-Context Learning and Out-of-Context Hallucination

Abstract

Large language models can infer and follow specific rules through in-context learning (ICL), yet may fail to remain faithful to provided evidence, giving rise to out-of-context (OOC) hallucination. Existing work has largely investigated the underlying mechanisms of ICL and OOC hallucination along separate lines. However, this separation obscures the mechanistic basis for these seemingly opposing outcomes. In this work, we study this contrast through the unified view of functional roles in dynamic context. The key observation is that role information can be internally represented without reliably constraining output. We derive a first-order decomposition to characterize the role effect and identify two necessary conditions for role span contributions to become consequential for generation. Causal analysis further shows that role effects act through key token choices. These findings support a common mechanistic view of ICL and OOC hallucination: each functional role defines the pattern of output decisions over which role span contributions must remain effective, and differences in this organization account for the seemingly opposing outcomes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.