acceptodds
Under review as a conference paper at ICLR 2027

Inference-Time Search in Discrete Diffusion Language Models, Guided by Internal-State Signals

Abstract

Inference-time search for diffusion language models assumes that a partially decoded state reveals, early and cheaply, whether a trajectory is worth continuing, and that the model's own internal state is the natural place to read it. We test that assumption on three testbeds spanning four orders of magnitude — two masked-diffusion Sudoku models (0.54M and 6.6M parameters) with an exact symbolic oracle, and LLaDA-8B on GSM8K — and find that internal-state signals track the model's belief, not the partial state's determinacy. We make this precise by splitting the variance of prefix value into a between-prompt part, which only allocation can use, and a within-prompt part, which is all that selection can use. Pooled AUROC, the standard metric, is dominated by the former and so measures difficulty: a difficulty-only signal scores 0.78 pooled and 0.50 within-prompt. Even within-prompt AUROC fails to price a selector: on the small model, confidence ranks branches above chance (0.64), yet choosing its top branch cuts accuracy from 0.69 to 0.36. We introduce a cheap estimator of within-prompt variance from branched rollouts and show it is set by the sampler — zero at temperature 0.3, and at 0.9 half or more of the remaining outcome variance on LLaDA-8B by r=0.5, where a perfect selector would add 15 points. Model confidence recovers none of it at 8B; a task-structural oracle recovers most of it on Sudoku. We recommend reporting any verifier by its realized selection value, at the sampler's operating point.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.