acceptodds
Under review as a conference paper at ICLR 2027

Near-Linear Adaptivity Gaps in Parallel Masked Generation

Abstract

Masked generators reduce decoding depth by filling several positions at once, but each round must decide both what values to generate and which positions to fill together. Existing work shows that adapting this choice can help, but leaves open whether the generated values themselves are necessary for fast and accurate generation. We prove that they are, even when the full target distribution is known and every per-position prediction is exact. On a four-symbol target family where every pair of positions is independent, a decoder that uses earlier values samples exactly in three rounds, whereas every randomized schedule that ignores them needs nearly linear depth to attain a fixed small error. We further prove that any short, accurate decoder must make its batch choices reflect the information revealed by those values. In contrast, for uniform sampling without replacement, generated values cannot improve the optimal schedule. These results identify generated values as a necessary scheduling resource, rather than merely a useful signal, and show when this necessity disappears.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.