acceptodds
Under review as a conference paper at ICLR 2027

When to Stop and Which to Return: Trajectory-Level Minimum Bayes Risk Decoding for Diffusion Language Models

Abstract

Masked diffusion language models can produce high-quality intermediate responses before denoising finishes, yet continued decoding may incur unnecessary computation or overwrite better responses. Existing approaches address these challenges separately. Early-commitment methods reduce computation but cannot recover earlier responses. Temporal voting exploits intermediate responses to improve quality but retains the cost of full decoding. We propose *Trajectory-MBR*, a training-free method that jointly decides when to stop and which current or earlier response to return. It reuses existing denoiser outputs to construct complete response candidates and performs trajectory-score-weighted minimum Bayes risk selection within a sliding window. Once the minimum weighted risk falls to or below a threshold, decoding stops and returns the selected response. We show that accurate trajectory-based utility estimates yield near-optimal selection within the window and provide an offline condition for improvement over full decoding. Across language and multimodal tasks, Trajectory-MBR maintains or improves average response quality with average denoiser-evaluation speedups of up to , while its token-level extension preserves response quality at a speedup on HumanEval.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.