acceptodds
Under review as a conference paper at ICLR 2027

Learning to Refine Partial Solutions with Recursive Discrete Denoising

Abstract

Structured reasoning requires predictions whose parts jointly satisfy constraints or an inferred rule. Reusing a small neural network across computation steps offers a parameter-efficient approach, but each intermediate answer can contain both useful assignments and errors. Learning to refine such answers is challenging when supervision provides completed solutions without intermediate reasoning trajectories. We propose the Discrete Reasoning Diffusion Model (DRDM), which trains a looped Transformer to reconstruct solutions from partially correct candidates. Masking and token replacement create completion and correction tasks while preserving the observed problem. A shared core performs several latent updates before predicting solution tokens; progressive decoding adds assignments, allows revision, and carries latent memory between steps. Across three training seeds, the approximately 7M-parameter model achieves 85.2% exact accuracy on Sudoku-Extreme, 87.3% on Maze-Hard, and 50.5% pass@2 on ARC-AGI-1, exceeding self-attention TRM by 9.7, 2.7, and 7.3 percentage points, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.