acceptodds
Under review as a conference paper at ICLR 2027

G2ARL: GROUP GFLOWNETS FOR AGENTIC REINFORCEMENT LEARNING

Abstract

Many agentic tasks admit multiple successful trajectories, yet reward maximization does not explicitly preserve their diversity. GFlowNets match a reward-proportional trajectory distribution, but a sparse reward over long interactions makes such a distribution difficult to learn. We introduce Group GFlowNet for Agentic RL (G²ARL), which combines group level flow reallocation with coverage-guided dynamic rebalancing for multi-turn language model agents. The flow reallocation loss learns relative probability allocation from pairwise trajectory comparisons, while dynamic gating and submodular replay retain verified successes selected for coverage and return them to groups that have none. Across ALFWorld, WebShop, and ScienceWorld, G²ARL achieves the highest success among all retrained baselines in every main benchmark setting. Across benchmark settings, G²ARL achieves 4.2 to 20.7 percentage points higher success than the strongest retrained comparator, while also producing shorter solutions with fewer repeated actions and loops.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.