When Does Mixing Candidate Sources Beat Resampling in Diffusion Language Models?
Abstract
Masked diffusion language models can spend extra generator calls in three ways: resample the same decoder, decode with a different rule, or query a different model. We ask when drawing from several of these sources beats resampling one. The analysis starts from a simple observation: repeated draws solve the prompts a source already handles, so the prompts left open concentrate on its weaknesses. For conditionally independent draws we derive how this residual set evolves and obtain a switching rule: another source gains ground when it succeeds where the current source fails, and a learned verifier returns only part of what it adds. The rule yields a greedy source order for tasks with an exact checker and a gated escalation rule for learned verifiers, both fixed on development data and tested prospectively. On Countdown and CRUXEval-I, mixing decoders and models outperforms resampling the released dVoting decoder by 16.8 to 21.2 points at the same budget, in the direction predicted, some smaller than predicted, and matches or beats the strongest single model given the same number of draws. With learned verifiers on MATH the gain shrinks to 3.1 to 5.5 points, and on GSM8K, where few prompts remain unsolved, it vanishes, as the analysis predicts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.