Iterative Value-Guided Tilting for Fine-Tuning Discrete Diffusion Models
Abstract
Discrete diffusion models have demonstrated considerable success across generative tasks ranging from natural language to biological sequences. However, fine-tuning these models on fine-grained black-box rewards, such as binding affinity or molecular property scores, remains challenging. We propose a novel algorithm for discrete diffusion models that requires only black-box reward evaluations and iteratively alternates between learning a soft-value function from samples of the current model and updating the generative model using the resulting value estimates. This iterative procedure progressively steers the model toward higher-reward distributions without requiring reward gradients or a pre-collected reward-labeled dataset. We validate our method on DNA sequence design and protein inverse folding tasks, outperforming comparable methods that rely only on black-box reward evaluations while remaining competitive with approaches that additionally exploit reward gradients or labeled data. Using a simple two-dimensional example, we further illustrate the benefit of successive rounds of value estimation and model optimization over a single optimization round.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.