Schrödinger’s Thinking: Reasoning via Superposition in Diffusion Large Language Models
Abstract
Diffusion large language models offer a promising alternative to sequential reasoning by predicting many parts of a solution in parallel. However, we find that their reasoning is still limited by an important form of sequential commitment: each prediction must collapse to a single discrete token, even when the future remains uncertain. We introduce reasoning via superposition, a framework that delays this commitment. Instead of selecting one token immediately, our approach represents an intermediate prediction as a continuous mixture of several plausible tokens. This allows multiple possible reasoning paths to remain active and jointly influence later predictions until the model has enough context to make a discrete decision. We further introduce reinforcement learning that teaches diffusion language models to reason through these superposed states. Across several diffusion language models and reasoning benchmarks, superposition improves reasoning accuracy at inference time without changing model parameters, and reinforcement learning further increases the gains to 7.3 points. Our analysis shows that preserving alternative hypotheses is especially important early in reasoning, when the correct future has not yet been determined.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.