ReStrike: Preventing Support Collapse in On-Policy GFlowNets
Abstract
GFlowNets are trained to sample objects in proportion to a reward. We show that the standard on-policy balance objectives have a silent failure: the sampler can become exactly calibrated on part of the space while putting no mass on the rest. There the on-policy loss and its gradient vanish, so the collapsed sampler is a global minimum tied with the correct one, and it can even score a lower held-out loss than a healthy model. On Bayesian network structure learning with five variables, where the posterior over all 29,281 graphs is computed exactly, plain sub-trajectory balance without exploration ends collapsed on 17 of 20 runs across two posteriors, and trajectory balance, modified detailed balance and local search share the trap, which is metastable and, on one posterior, batch-dependent. On a fragment-based molecular benchmark the loss again prefers the sampler with the fewest modes, on two of three landscapes. The standard ε-uniform exploration floor prevents the collapse by the end of training but not the calibration error, even at four times the budget. ReStrike, which replays the distinct objects the run has found, drawn by a softmax over log-reward, recovers both posteriors at the same budget of fresh trajectories where floors, prioritized replay, a published floor-plus-replay schedule and local search fall short, and finds an order of magnitude more modes than the floor on a tiered grid. We prove that replay from inside a collapsed support cannot move it, and on one eight-node posterior replay can freeze where that schedule does better. Finally, the released code of prioritized replay training inverts its own prioritization on tiered rewards.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.