acceptodds
Under review as a conference paper at ICLR 2027

The Unembedding Bridge: From Hidden-State Movement to In-Context Learning Predictions

Abstract

Mechanistic accounts of in-context learning (ICL) characterize internal representations and causal components, whereas output-level accounts measure demonstration-induced changes in the predictive distribution. We argue that the unembedding bridges these levels: ICL reflects the interaction of a context-induced hidden-state movement with training-shaped output geometry. Across seven open-weight language models and seven classification tasks, ICL final states generally have higher mean unembedding alignment than paired zero-shot states, against a frequency-organized, spectrally concentrated output geometry. Synthetic-model training, synthetic–LLM comparison, and theory show how cross-entropy attraction and repulsion can create this geometry when the optimizer rescales rows individually. Splitting the movement at the label-contrast space of each task, we find that its label part lowers mean alignment, because the demonstrated labels point against the mean unembedding direction, whereas its orthogonal part raises it and carries the vocabulary-wide shift; remapping the labels to non-outlier symbols removes the label part but leaves the orthogonal shift. Finally, targeted ablation, attribution, probability shifts, and steering identify induction heads as a principal causal source of the movement. Thus training supplies the geometry, demonstrations the movement, and their interaction the behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.