acceptodds
Under review as a conference paper at ICLR 2027

Multiscale Predictive Dynamics Make Model Reasoning Observable and Controllable

Abstract

Language models routinely perform long multi-step reasoning, yet they sometimes loop without stopping, their chains of thought are not a reliable record of what they compute, and we lack an account of how local computations are coordinated into reasoning trajectories. Here we procedurally generate SchemaBench, multi-phase schema tasks that vary roles, temporal structure, and complexity and allow ground-truth parsing of reasoning traces. We evaluate 10 LLMs' behavior and the residual-stream activations of Phi4-mini-R and Qwen3-4B. Their latent computation is organized at three levels: a) task-general operations conduct local computations with a shallow grammar, b) schema-level representations of roles, phases, and progress bias which operations follow, c) slower control states track the switching and completion of reasoning. Across levels, reasoning is organized in a low-dimensional subspace of the residual stream, where a state-dependent vector field predicts the next latent transition and reasoning 5-10 steps ahead. Causally injecting a STOP control-vector into Phi4-mini-R ends traces that could otherwise loop to the token cap, shortens response length by 4000+ tokens without sacrificing accuracy, and transfers to an unseen mathematical task. LLM reasoning is thus organized as multiscale predictive dynamics within a shared low-dimensional field, which makes hidden reasoning observable and controllable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.