acceptodds
Under review as a conference paper at ICLR 2027

Gentoo: Draft Continuously, Decode Discretely

Abstract

Masked diffusion models speed up generation by committing multiple tokens per round, each sampled independently from its predicted distribution. Tokens committed together can therefore conflict, especially with fewer rounds. To address these conflicts, some methods select which tokens to commit together or jointly denoise auxiliary continuous variables with the tokens, but they still settle instance content during token commitment. Our key insight is therefore to draft continuously and decode discretely: a fine-grained, well-structured draft gives parallel token predictions a common reference before the first commitment. Gentoo first uses a pretrained latent generator to draft an instance in representation space. A learned adapter then provides this draft as shared context to a pretrained masked diffusion decoder, which translates it into tokens and progressively completes the sequence. To align training encodings of real samples with drafts generated at inference, we partially noise each encoding and let the latent generator denoise it, while keeping the sample's own tokens as targets. With the shared draft, tokens committed in the same round become more consistent, and quality depends far less on the number of rounds. Gentoo improves generation quality over MaskBit and Meissonic while using only 62.5% and 64% of their respective original NFE budgets, counting both continuous and discrete stages. These configurations lower inference time by 16% on ImageNet and 53% in text-to-image generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.