Mixture of GFlowNets
Abstract
Generative Flow Networks (GFlowNets) learn a policy that samples compositional objects in proportion to an unnormalized reward . This makes them useful wherever many diverse, high-reward candidates are needed and oracle queries are expensive, as in AI for Science. In practice, a single GFlowNet often fails to cover all modes of a complex distribution, and the obvious ways to scale compute saturate: larger batches and data-parallel gradient averaging draw every trajectory from one, possibly biased, policy, so extra rollouts rediscover known modes instead of finding new ones. We explore whether this mode-bias problem (mode collapse in the extreme) can be addressed with a population of asynchronously trained policies which divide the state space between them. **Mixture of GFlowNets (MoG)** replaces the single forward policy with a _community_ of independent GFlowNets coupled through one jointly trained soft classifier, which assigns each member its own region of the state space through reward shaping. Mixing the members under their learned partition functions provably recovers the target . **MoG exchange (MoGX)** also lets the classifier route sampled trajectories: a trajectory the classifier assigns to another member is sent to that member and trained on there, with no importance sampling. We evaluate both methods by keeping the environment-interaction budget fixed while scaling the community size, turning parallel _compute_ into mode _coverage_ without extra oracle calls. On HyperGrid, across six reward landscapes and communities of up to , MoG and MoGX keep finding new modes as grows, while data-parallel training saturates. On a genomic ideotype-design task, MoGX covers % of the optima at , against % for independent agents and % for data-parallel training. MoG also has a lower per-iteration synchronization cost than data-parallel training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.