Mean-to-Score discrete diffusion: posterior-mean denoisers for score entropy
Abstract
Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios. Although positivity ensures nonnegative reverse jump rates, it does not ensure Bayes realizability: at a fixed noisy state, all candidate score ratios must arise jointly from a single clean-token posterior under the forward kernel. We show that the score-entropy loss has the correct population optimum but does not enforce Bayes realizability away from it. In a trained pure-uniform SEDD checkpoint, roughly one quarter of complete score vectors violate the coordinate box, while more than half satisfy every coordinate bound but remain materially incompatible with any valid clean-token posterior. Although the corresponding continuous-time reverse jump rates remain nonnegative, these violations can induce negative pre-normalization weights in the finite-step sampler update. Projecting the checkpoint's raw scores onto the bridge polytope removes all observed negative weights and lowers external generative PPL from to without changing the sampler. To enforce Bayes realizability by construction rather than through post-hoc projection, we introduce mean-to-score (M2S): the network predicts a clean-token posterior mean and converts it to the score through an exact kernel-dependent linear map. The map applies to any known coordinate-wise continuous-time Markov chain (CTMC) satisfying a mild support condition. For uniform corruption, it maps the probability simplex onto the bridge polytope; for absorbing-mask corruption, the resulting objective recovers MD4 exactly. In a controlled 28.4M-parameter CIFAR-10 comparison, M2S lowers test BPD from to and FID-50k from to . A 170M-parameter M2S model trained on approximately 131B OpenWebText token slots outperforms the evaluated pure-uniform SEDD, GIDD, and Neural CTMC checkpoints at every tested sampling budget, reaching mean generative PPL at 128 steps compared with for the strongest pure-uniform baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.