Keeping the Reader's Place: Why Revealing Evidence One Step at a Time Helps Masked Diffusion Language
Abstract
A masked diffusion language model rereads its prompt on every forward pass, so what the prompt shows can change while the answer is being written. Revealing one record at a time along a chain of pointers turns a hard multi-hop prompt into a sequence of easy local reads, and it is tempting to credit the gain to a model that has learned where to look. We take the gain apart on tasks whose dependency structure is explicit, holding the model, the items, and the forward passes fixed while varying what is visible when, what the reader was trained on, and whether the order of reads is decided during generation or before it. A LLaDA-8B model fine-tuned to read one record at a time solves chains longer than any it saw in training. The same model fine-tuned on full prompts and handed exactly the same records all at once falls well short, and its errors have a consistent signature: when a value recurs in the chain, it resumes from the wrong occurrence, far more often than chance predicts. Revealing one record at a time removes this bookkeeping from the reader. The order that does so need not come from the model: a text parser computes it from the prompt before generation begins. Two further controls locate what remains. The strength of a static baseline depends on what it was trained on, and when neighboring answers are independent, as on chains of IMDB reviews, classifying each record on its own does at least as well as reading the chain on average. Revealing evidence one step at a time helps a diffusion reader because it keeps the reader's place, and the schedule that does so can be fixed in advance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.