acceptodds
Under review as a conference paper at ICLR 2027

Rethinking GFlowNet Training: Negative Reinforce Dynamics

Abstract

Generative Flow Networks (GFlowNets) are models that learn to sample discrete objects proportional to a target reward. They are trained by minimizing balance errors over sampled trajectories, which inherently ties correctness to exploration: if a region is not visited, balance cannot be enforced there. Adaptive exploration is therefore crucial, yet standard training relies on a semi-gradient update that treats the sampling distribution as fixed at each optimization step. On the other hand, the full gradient reveals a second mechanism that reduces the observed errors in an anti-exploratory manner. In this paper, we establish a connection between the full gradient and min-min dynamics between the GFlowNet and a decoupled sampler, showing that the sampler reduces observed violations by avoiding high-error regions. This insight motivates a bridge between two-player games and single-policy dynamics, which underpins the design of our proposed method, Negative Reinforce Dynamics (NRD). By tying the parameters of an adversarial game together, we recover a single-policy dynamic that extends exploration while maintaining correctness. In standard benchmarks, our method outperforms existing approaches in both convergence speed and exploration efficiency, without training a separate teacher or auxiliary exploration policy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.