acceptodds
Under review as a conference paper at ICLR 2027

A Family of Control Decisions Bypasses the J-Space

Abstract

Reasoning language models carry reportable content in a small, causally privileged subspace of the residual stream, a verbalizable “global workspace” built from that content Gurnee2026. We ask whether this content-derived workspace also carries the model's control decisions: choices like when to stop, call a tool, or refuse. Clamping the workspace to its clean value removes most of the causal effect of a content perturbation that would otherwise pass through it, but leaves a control decision's effect essentially untouched. This holds under the actual sparse-gradient-pursuit decomposition, beyond a linear approximation, and generalizes across model architectures and across a family of control tokens including refusal. Content is broadcast through the workspace; control decisions are computed outside it and surface only through their own readout. The implication we note is that a model can therefore report a control decision faithfully without ever pointing to what caused it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.