Deterministic Joint-Risk Sampling for Efficient Discrete Diffusion
Abstract
Masked diffusion language models repeatedly refine masked tokens, often spending forward passes on predictions that have already stabilized. Fixed transfer budgets do not adapt the number of commitments to observed token stability, creating a trade-off between unnecessary refinement and premature commitment. We introduce Deterministic Joint-Risk Sampling (DJRS), which ranks masked tokens using confidence, margin, Jensen–Shannon divergence, and prediction change, then admits additional low-risk commitments under a progress-dependent budget. On LLaDA-8B-Instruct, DJRS reduces the number of forward evaluations (NFE) by 29% on GSM8K at 64 steps (64.0 to 45.6) with a 0.6-percentage-point accuracy decrease, and raises MATH-500 accuracy from 27.8% to 31.4% while reducing NFE by 22%. In a separate Dream-v0-Instruct-7B HumanEval evaluation, DJRS reaches 38.4% pass@1 at 39.8 NFE, compared with 32.9% at 64.0 NFE for Default. Trace analysis shows that margin most consistently separates stable from flip-prone masked positions, while JSD shows less monotonic associations with flips. We present DJRS as a practical adaptive-decoding heuristic rather than a calibrated risk guarantee.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.