: Stratified Scaling Search for Test-Time Scaling in Diffusion Language Models
Abstract
Test-time scaling asks how a fixed diffusion language model (DLM) can use additional inference compute to improve generation without retraining. Standard Best-of- sampling spends this compute on generating independent complete outputs and selects only after decoding, leaving the generation process itself unchanged. We introduce (Stratified Scaling Search), which instead allocates additional compute during denoising. At each denoising step, expands multiple partial trajectories, evaluates their one-step clean predictions with a black-box verifier, and resamples promising trajectories while preserving diversity. We derive this procedure from a KL-regularized reward-tilted target distribution and use incremental child-parent potentials whose product telescopes to the desired terminal reward tilt, making the ideal search objective invariant to the number of denoising steps. Our default verifier is training-free and requires neither a learned reward model nor ground-truth labels. Across six benchmarks, consistently improves over standard diffusion decoding at every tested inference budget. Its performance relative to other test-time scaling methods depends on the reliability of intermediate predictions: informative intermediate signals favor search during denoising, whereas costly or unreliable process-reward-model scoring can favor output-level selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.