acceptodds
Under review as a conference paper at ICLR 2027

Generation-Free Verification in Diffusion Language Models via Anchor Chain Scanning

Abstract

As language models solve increasingly complex reasoning problems, the need for reliable verification grows accordingly. Notably, a single error can propagate through a reasoning trajectory and yield a plausible yet incorrect answer. To address this challenge, recent work has explored generative verifiers based on autoregressive language models and diffusion language models (DLMs). However, such generative methods can hallucinate during verification and often waste computation on unnecessary parts of the reasoning trajectory. To overcome this, we propose De-Anchoring: a generation-free DLM-based verifier that focuses on information anchor chains, defined as repeated occurrences of key information throughout a reasoning trajectory. De-Anchoring simultaneously masks all occurrences of each anchor and uses their DLM reconstruction probabilities to measure how strongly the remaining reasoning supports the original information. Across practical verification tasks, including Best-of-N selection, unsupervised error diagnosis, and auxiliary reward signals for reinforcement learning, De-Anchoring substantially outperforms generative verification baselines. Moreover, De-Anchoring remains robust under aggressive cost reduction, achieving comparable performance with only 12.5% of the generative verification cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.