Learning the Möbius Attractor: From Gradient Descent to Reasoning by Superposition
Abstract
LLM Reasoning often requires keeping several possibilities until later evidence shows which ones matter. Through superposition reasoning, a Transformer can do this by representing a set of candidate states in one hidden state and updating the whole set over several steps. Prior work shows that carefully chosen parameters can represent it, but representation alone does not explain whether gradient training can learn it. We give sufficient conditions under which training learns the full sequence of target intermediate states, rather than only the final answer, and generalizes to new tasks from the same distribution. The conditions have a simple meaning: the model must preserve the current candidate set and be able to produce the next one, while the training signal must directly correct inaccurate intermediate states. We then study a practical problem. When intermediate-state labels are limited, which steps should be supervised? For a stated model of how supervision reduces error, we derive the optimal continuous allocation of a fixed label budget. The solution gives more labels to steps where one label is expected to reduce more error at lower cost, and it motivates a practical rule based on measured improvement. Together, these results explain when reasoning by superposition is learnable and how limited supervision should be distributed across its intermediate states.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.