Rethinking Reasoning Through Masks
Abstract
Reasoning is built from intermediate results: each step of a proof, program or iterative computation rests on what earlier steps established. Masked diffusion models (MDMs) generate by repeatedly filling in blanks, in any order and with the option to revise, and this flexibility has made them a promising alternative to left-to-right language models for reasoning. Yet their gains have fallen short, and it is unclear whether better decoding orders, larger models or stronger revision can close the gap. Here we provably show that none of them removes the fundamental bottleneck: between rounds, an MDM remembers only its tokens, and a token must be either an answer or a blank. Without revision, an MDM giving a short answer is effectively one forward pass deep, however many rounds it runs and however large it grows. Revision lets rounds build on one another, but only by paying for extra rounds or scratch tokens that grow with input length. And when the intermediate results a task needs are not yet answers, no amount of output space can hold them. Passing hidden states from one round to the next removes these limits. This Latent-State MDM, trained on at most 40 steps of the Rule 110 cellular automaton, executes 96 steps without error, where GPT-6 Astra solves only 6 of 10 cases, and lifts Sudoku-Extreme accuracy from 8.95% to 73.95% on the same 6.6M-parameter backbone, surpassing MDMs 46 times larger. Reasoning requires unfinished computation to survive before it becomes an answer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.