Identifying Token and Task-Relevant Representational Subspaces in Large Language Models
Abstract
Understanding how language models predict particular tokens and perform specific tasks requires identifying representational subspaces that selectively support these behaviors across contexts. We develop a unified framework based on gradients of token scores and behavioral losses with respect to hidden states. Averaging gradient outer products yields a sensitivity matrix whose leading eigenvectors capture directions of greatest local sensitivity. At the token level, we separate the context-averaged gradient direction from a residual subspace that captures context-dependent sensitivity orthogonal to it. We extend this framework to task selectivity by contrasting sensitivity matrices for target and comparison tasks, yielding Contrastive Task Directions. We find that token-relevant residual subspaces in Qwen and Llama generalize across independent contexts, with sensitivity increasingly concentrated in later layers. On held-out multihop reasoning prompts, ablating subspaces associated with an intermediate concept selectively reduces the probability of the correct downstream answer, while counterfactual interventions can redirect generation toward the answer implied by a substituted concept. Extending this analysis to task-level behavior, we find that ablating contrastive subspaces across seven Qwen and Llama models produces selective effects on syntactic, world-knowledge, and commonsense-reasoning benchmarks, with less off-target disruption on average than non-contrastive subspaces. Together, these results connect local gradient geometry to selective causal effects at both token and task scales.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.