Geometry-guided Exploration with Adaptive Termination for GFlowNets
Abstract
Generative Flow Networks (GFlowNets) learn to sample compositional objects in proportion to their rewards, requiring exploration that discovers diverse high-reward objects. Existing exploratory policies can repeatedly acquire experiences that remain useful for learning but are already available through replay. We propose *Geometry-guided Exploration with Adaptive Termination (GEAT)*, a simple yet effective algorithm for efficient exploration in GFlowNets. Its core mechanism, *novelty-aware termination (NAT)*, generates terminal candidates through local backtracking and reconstruction, then selects the candidate with the highest novelty relative to reference states sampled from replay. Only the selected candidate requires a task-reward evaluation. To guide this selection, we derive *geometric novelty* from the derivative of geometry-aware Shannon entropy. The resulting score incorporates similarities between state representations and ranks candidates by their first-order effect on the entropy of the reference distribution. GEAT uses this score both as an intrinsic reward for the exploratory policy and as the terminal-selection criterion. Across extensive benchmarks, we show that GEAT consistently discovers high-reward modes faster and achieves broader mode coverage while maintaining competitive reward quality and distributional accuracy, demonstrating improved sample efficiency in mode discovery across diverse tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.