Content Before Position: A Bayesian Decision Framework for Diffusion Language Model Decoding
Abstract
Diffusion language models predict tokens across a canvas in parallel, yet decoding typically prioritizes positions by local confidence. This can overlook collective evidence for a token distributed across several candidate positions. We introduce content-before-position decoding, a principle that separates what to generate from where to commit it. A one-step Bayesian formulation connects maximum confidence position-first decisions to an aggregate-evidence, token-first endpoint. Building on this framework, we propose COLA (Collective-evidence Ordering and Local Acceptance), a training-free decoder that uses token-first proposals and validates placement locally. It coordinates parallel allocations from each model prediction and regulates commitment and revision across denoising steps. On DiffusionGemma, COLA modestly improves on the already high PARALLELBENCH quality of the model's default Entropy Bound (EB) decoder, while reducing model evaluations per generated token by 39.4% and latency by 33.2%. Across six downstream tasks, aggregate model evaluations per generated token fall by 73.7% relative to EB, with higher accuracy on both mathematics benchmarks and pass@1 gains of 20.7 and 10.5 percentage points on MBPP and HumanEval, respectively. These results establish content selection as a practical design axis for efficient parallel diffusion decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.