Towards Value-Aware Unmasking: Local Action Value in Diffusion Language Models
Abstract
Diffusion language models (DLMs) must decide which masked position to commit next, yet existing endpoint predictors estimate trajectory-level success, and lookahead methods compare candidate futures. We ask whether the frozen representation already encodes value differences between alternative commitments. We introduce Local Action-Value Readout (LAVR), a framework that formalizes policy-relative action value for unmasking decisions, estimates it through matched forced-and-reference rollout branches with candidate centering, and reads it from frozen hidden states against output-and-trajectory controls. In an 800-prompt GSM8K study on Nemotron-Diffusion-3B, the hidden-state readout improves global state-value prediction over the output-statistics baseline (R² = 0.4812 vs. 0.3380) and supports selective generation. For local decisions, candidate-specific path log-likelihood (Path-LL) value replicates on a disjoint 600-document pool and improves pairwise ranking on an independent 400-document study across all six model–policy/target conditions. By contrast, task-correctness action gains are not reliably positive, and frozen transfer decreases concordance. This contrast shows that readable local continuation value does not automatically imply task-correctness value, making reward alignment central to task-aligned unmasking. Code and evaluation artifacts are available in an anonymous repository.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.