Data first, Rewards Later: GFlowNets Across Pre-Training and Post-Training
Abstract
Discrete diffusion models and Generative Flow Networks (GFlowNets) are two seemingly alternative approaches to modelling discrete data. They differ in terminology and parametrisation, but most critically in how they are trained. While discrete diffusion models are primarily trained from data, GFlowNets are primarily trained from rewards. We show that this separation is not fundamental: both model classes admit a common flow-score formulation that allows training objectives to transfer between them. Building on this connection, we propose a pre-training/post-training regime analogous to the typical LLM training pipeline. A GFlowNet can first be pre-trained directly from data, without an explicit reward model, and subsequently post-trained with standard reward-based GFlowNet objectives without changing the underlying generative model. This combines the flexible construction spaces and structured generation native to GFlowNets with efficient learning from existing data characteristic of discrete diffusion models. On peptide generation, data-based pretraining substantially improves subsequent reward-query efficiency under identical post-training, reaching 88.6% high-reward samples compared with 73.5% when training from untrained initialization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.