acceptodds
Under review as a conference paper at ICLR 2027

Data first, Rewards Later: GFlowNets Across Pre-Training and Post-Training

Abstract

Discrete diffusion models and Generative Flow Networks (GFlowNets) are two seemingly alternative approaches to modelling discrete data. They differ in terminology and parametrisation, but most critically in how they are trained. While discrete diffusion models are primarily trained from data, GFlowNets are primarily trained from rewards. We show that this separation is not fundamental: both model classes admit a common flow-score formulation that allows training objectives to transfer between them. Building on this connection, we propose a pre-training/post-training regime analogous to the typical LLM training pipeline. A GFlowNet can first be pre-trained directly from data, without an explicit reward model, and subsequently post-trained with standard reward-based GFlowNet objectives without changing the underlying generative model. This combines the flexible construction spaces and structured generation native to GFlowNets with efficient learning from existing data characteristic of discrete diffusion models. On peptide generation, data-based pretraining substantially improves subsequent reward-query efficiency under identical post-training, reaching 88.6% high-reward samples compared with 73.5% when training from untrained initialization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.