acceptodds
Under review as a conference paper at ICLR 2027

Rational Manifolds: The Spatial Organisation and Control of LLM Reasoning

Abstract

The thinking traces of large language models (LLMs) interleave behaviours such as planning, exploring and verifying, yet little is known about how these behaviours are arranged in latent space. Linear representations provide a framework to explain and control model behaviour but the directions it relies on are often unstable. We provide a mechanistic recount of LLM reasoning by interpreting regions instead of directions as the semantic units of latent space, partitioning activations via a Mixture of Factor Analyzers (MFA). Our experiments reveal that these regions correspond to recognisable problem-solving behaviours with complex nonlinear boundaries and whose local concept directions consistently disagree with global ones. Causal interventions show that regions offer limited control over individual behaviours but can have a strong impact on reasoning performance: diverting traces away from a small subset of load-bearing regions critically degrades accuracy, but carefully redirecting the path a model takes through specific regions can also make reasoning shorter and more accurate. This mechanism is found to differ geometrically from linear steering. Together, these findings suggest that localised descriptors, such as latent space regions and the paths through them, can inform new methods for both interpretation and control of reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.