Balancing Quality and Coverage in Diffusion Language Models with Adaptive Risk Portfolios
Abstract
Generating multiple candidate responses can improve reasoning performance, but its effectiveness depends on balancing individual response quality, solution coverage, and computation. In diffusion language models, conservative token commitment can improve individual responses while limiting the coverage gained from additional samples. We introduce Confidence-Adaptive Risk Portfolios (CARP), a training-free decoding method that allocates complementary commitment policies across a candidate population. A budget-dependent reserve of trajectories uses an adaptive rule that constrains the cumulative confidence deficit of eligible token proposals at each denoising step, while retaining stochastic selection within the eligible set. The remaining trajectories follow two fixed tempered confidence-threshold policies to support broader exploration. A progress-based computation guard limits prolonged constrained decoding. Together, these components balance conservative commitment and stochastic exploration without additional training or correctness feedback. Experiments on MATH-500, HumanEval, MBPP, and GSM8K with LLaDA and Dream compare CARP with TCT and ODD. CARP achieves the highest mean pass@ at the reported budgets on MATH-500 and HumanEval, improves early-budget performance on MBPP, and remains competitive as GSM8K approaches saturation. These results support population-level commitment allocation as a practical approach to improving response quality and solution coverage in diffusion language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.