acceptodds
Under review as a conference paper at ICLR 2027

A Sparse Prototype Readout for LLM In-Context Classification

Abstract

What class-level information in a language model’s contextual representations drives in-context classification? How is that information transformed into a prediction? Across four open-weight LLMs, we trace this computation from contextual representations to the final readout. Contextual processing progressively constructs a low-dimensional class geometry whose class-mean directions are causally linked to the model’s decision: continuously exchanging these directions reverses the output. We then follow this decision variable into a sparse set of attention heads. Contextual class means are reflected in class-mean keys, yielding headwise query–prototype margins in the native attention computation. Under the same geometric intervention, these margins vary approximately linearly and cross near zero where the contextual class-mean directions meet. On held-out instances, a fixed sparse combination of the native prototype margins predicts the model’s output with to and % to % prompt-level agreement. Thus, we identify a class-level geometric decision variable for in-context classification and a sparse prototype-based circuit that reads it out.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.