acceptodds
Under review as a conference paper at ICLR 2027

Beyond Blind Generation Scaling: Verification-Guided Inference for Diffusion Language Models

Abstract

Test-time scaling typically improves language-model performance by spending additional inference compute on generation, implicitly assuming that more generation compute directly yields more reliable outputs. We show that this view is incomplete for diffusion language models. By decomposing model outcomes according to their observability, we identify a compute-dependent loud-to-silent-to-correct transition in the output distribution: as denoising compute increases, loud failures decline before semantic errors disappear, creating an intermediate regime in which surface-valid but incorrect outputs become dominant among the remaining failures before correctness eventually catches up. This reveals that surface validity can improve before semantic correctness. Motivated by this failure transition, we introduce *verification-guided inference*, a new test-time inference paradigm that moves beyond blind generation scaling by separating candidate generation from output assurance. The model first generates with a moderate budget, filters directly observable validity failures with a lightweight deterministic check, and applies an independent verifier to the remaining candidates; only rejected candidates receive higher-budget regeneration. Across code generation and mathematical reasoning, the transition persists across diffusion model families and evaluation controls. On security-sensitive programming tasks, silent failures can additionally include functionally valid yet vulnerable programs. Independent verification detects 72–83% of silent failures with a single pass costing only 2.3% of the token-forward compute of 256-step generation. Under full coverage, verification-guided inference improves correctness from 41.1% to 48.0%, reduces silent failures from 45.8% to 34.0%, and lowers average generation-equivalent compute from 256 to 243.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.