acceptodds
Under review as a conference paper at ICLR 2027

Choose Before You Decode: Latent Q-Guided Generation for Diffusion Language Models

Abstract

Continuous latent diffusion language models generate an entire response by iteratively refining a continuous latent representation, yet existing samplers typically assign the same denoising budget to every trajectory. Consequently, test-time methods such as Best-of- must fully denoise and often decode every candidate before selection. We introduce Latent Q-Guided Generation (LQG), a value-guided inference framework that evaluates latent trajectories before denoising is complete. LQG formulates the reverse process as a finite-horizon Markov decision process with terminal text utility and trains a Transformer critic through temporal-difference policy evaluation of the frozen denoiser. At inference time, multiple trajectories are advanced to an intermediate decision step, where the critic identifies promising candidates and prunes the remainder. Across unconditional and conditional generation tasks, LQG consistently improves expected terminal utility under independently specified scorers while substantially reducing denoising and decoding costs relative to terminal-scored Best-of-. Ablations further establish the importance of cross-trajectory selection and local margin regularization, while revealing an early-horizon limit on reliable trajectory ranking. These results demonstrate that latent value estimation provides a principled mechanism for efficient test-time generation in latent diffusion language models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.