acceptodds
Under review as a conference paper at ICLR 2027

IDRF: Inverse-Distilled Reward Fine-Tuning of Masked Discrete Diffusion Models

Abstract

Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized reward objective, IDRF replaces the intractable sequence-level KL penalty with an inverse-distillation regularizer. We prove that the population inverse-distillation loss upper-bounds the KL between complete sequences of any length. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and opti- mize reward with a clipped policy-gradient objective over the student’s denoising trajectories. We evaluate IDRF across biological, text, and image domains. In each domain, it reaches high reward with up to 32× fewer sampling steps and avoids the reward hacking that reward-only fine-tuning exhibits

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.