IDRF: Inverse-Distilled Reward Fine-Tuning of Masked Discrete Diffusion Models
Abstract
Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized reward objective, IDRF replaces the intractable sequence-level KL penalty with an inverse-distillation regularizer. We prove that the population inverse-distillation loss upper-bounds the KL between complete sequences of any length. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and opti- mize reward with a clipped policy-gradient objective over the student’s denoising trajectories. We evaluate IDRF across biological, text, and image domains. In each domain, it reaches high reward with up to 32× fewer sampling steps and avoids the reward hacking that reward-only fine-tuning exhibits
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.