Learning to Refine Partial Solutions with Recursive Discrete Denoising
Abstract
Structured reasoning requires predictions whose parts jointly satisfy constraints or an inferred rule. Reusing a small neural network across computation steps offers a parameter-efficient approach, but each intermediate answer can contain both useful assignments and errors. Learning to refine such answers is challenging when supervision provides completed solutions without intermediate reasoning trajectories. We propose the Discrete Reasoning Diffusion Model (DRDM), which trains a looped Transformer to reconstruct solutions from partially correct candidates. Masking and token replacement create completion and correction tasks while preserving the observed problem. A shared core performs several latent updates before predicting solution tokens; progressive decoding adds assignments, allows revision, and carries latent memory between steps. Across three training seeds, the approximately 7M-parameter model achieves 85.2% exact accuracy on Sudoku-Extreme, 87.3% on Maze-Hard, and 50.5% pass@2 on ARC-AGI-1, exceeding self-attention TRM by 9.7, 2.7, and 7.3 percentage points, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.