In-Context Learning Operates as Concept Subspace Learning
Abstract
Regression and Bayesian accounts of in-context learning (ICL) explain how demonstrations can induce predictors, while mechanistic analyses often identify compact activation directions that steer prompted behavior. However, it remains unclear whether structured demonstrations induce low-dimensional concept inference. We study this question through a concept-subspace view of ICL, in which tasks vary only along intrinsic concept coordinates, although inputs are observed in a high-dimensional ambient space. For ridge and least-squares ICL proxies, prediction decomposes exactly into concept-coordinate regression and off-subspace leakage. We show that an idealized single-layer linear attention head optimal for this task family instead stores the concept subspace in its weights and reads the prompt only through it, so patching its written subspace carries the head's entire pooled clean–corrupted effect. Testing this prediction in Llama-3-8B on CounterFact-derived multi-relation prompts, we find that a -dimensional subspace of the -dimensional residual stream restores of the clean–corrupted accuracy gap, whereas patching the complementary subspace restores . Concept swaps redirect predictions toward injected relations, while random and cross-task matched-rank controls are largely ineffective. Additional experiments on Qwen2.5-7B, a controlled cross-lingual rule task, context-defined nonce mappings and two-step compositional rules show the same qualitative pattern. These results support concept subspaces as compact, task-aligned mediators of recoverable ICL behavior in structured task families, without implying full-circuit recovery. Our code is available at the following anonymous repository https://anonymous.4open.science/r/ICL-Concept-Subspace.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.