Logical subspace in LLMs
Abstract
Recent work has identified a human brain network specialized for abstract formal reasoning (Kean et al., 2025). Does the same hold true in language models? To answer this question, we introduce the minimal viable subspace (MVS) method, which searches for the lowest-rank activation subspace at a layer that preserves task performance when retained alone. Using MVS, we demonstrate low-rank subspaces supporting logical inference on Gemma and Qwen models. Furthermore, these subspaces exhibit dissociation from other tasks: retaining late logic subspaces preserves inference while impairing factual knowledge, working memory, cognitive control, and arithmetic; ablating them reduces logical inference accuracy to chance while largely sparing these other capacities. Our results suggest a functionally-localized core machinery for logic akin to that in the human brain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.