Logical Equivariance in Diffusion Language Models
Abstract
Rewriting a problem without changing its answer can change whether a language model solves it, even when average accuracy barely moves. We study this as logical equivariance: predictions should transform consistently with answer-preserving input changes. In four diffusion language models, paired evaluations on Sudoku, code, and mathematics reveal many problems solved in one presentation but not another. We build a Sudoku benchmark whose forced targets transport exactly under rotations and reflections, and prove that exact grid equivariance leaves row-major relative-score attention only two or three offset classes, so stability must be learned. In post-training, Reasoning-State Masking (RSM) improves solving but alone reduces stability; presentation randomization (PR) removes this cost, and its gain in normalized reliability is 8.3 points larger with RSM. Matched controls reproduce this interaction at two scales and trace its stability gain to training on partial states at RSM's density. The combined recipe raises LLaDA-8B's all-view correctness from 47.7% to 86.1% alongside higher accuracy, and improves worst-view maze success by 7.3–18.8 points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.