acceptodds
Under review as a conference paper at ICLR 2027

SONAR: Spectral Diversity for Lost Reasoning Modes in Masked Diffusion LMs

Abstract

Effective exploration for complex reasoning problems remains a critical challenge in masked diffusion language models (MDLMs), where standard sampling frequently collapses into redundant failure modes. Raising the sampling temperature can increase diversity but also degrade solution quality, leaving the full reasoning capabilities of these models largely underutilized. To unlock these capabilities, existing approaches are primarily concerned with algorithms for token-reveal decisions, or promoting exploration over the vocabulary space. These methods, though effective to varying degrees, generally overlook the rich representation structure of the MDLMs. In this paper, we thus propose Sonar, a sampling method that first scans for overlap via a spectral objective in RKHS, subsequently transmitting the gradient of this objective throughout the model. Sonar then utilizes this gradient feedback to steer samples to systematically uncover lost reasoning modes of the MDLM. Across HumanEval (coding) and GSM8K (mathematics), Sonar improves fixed-budget reasoning success for LLaDA-8B-Instruct and Dream-7B-Instruct. It also expands solution coverage and reduces algorithmic collapse by uncovering distinct valid solution strategies missed by ordinary sampling, providing empirical evidence for a capability–sampling gap in MDLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.