When to Stop and Which to Return: Trajectory-Level Minimum Bayes Risk Decoding for Diffusion Language Models
Abstract
Masked diffusion language models can produce high-quality intermediate responses before denoising finishes, yet continued decoding may incur unnecessary computation or overwrite better responses. Existing approaches address these challenges separately. Early-commitment methods reduce computation but cannot recover earlier responses. Temporal voting exploits intermediate responses to improve quality but retains the cost of full decoding. We propose *Trajectory-MBR*, a training-free method that jointly decides when to stop and which current or earlier response to return. It reuses existing denoiser outputs to construct complete response candidates and performs trajectory-score-weighted minimum Bayes risk selection within a sliding window. Once the minimum weighted risk falls to or below a threshold, decoding stops and returns the selected response. We show that accurate trajectory-based utility estimates yield near-optimal selection within the window and provide an offline condition for improvement over full decoding. Across language and multimodal tasks, Trajectory-MBR maintains or improves average response quality with average denoiser-evaluation speedups of up to , while its token-level extension preserves response quality at a speedup on HumanEval.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.