Generate Multiple Actions, Execute One: Experience-Guided Action Selection for LLM Agent Reinforcement Learning
Abstract
Group-in-Group Policy Optimization (GiGPO) improves credit assignment in reinforcement learning for LLM agents by comparing the returns of actions already executed at the same state. Comparing more candidate actions could improve action selection and provide additional learning signals. However, directly obtaining terminal feedback for each candidate requires completing a separate downstream trajectory, which involves many additional environment interactions. To avoid this cost, we introduce Experience-Guided Action Selection (EGAS), a method that uses an experience pool of historical state–action outcomes to guide both action selection and policy learning. At the current state, EGAS uses the experience pool to estimate whether an alternative action could be more likely to lead to task success than the initially generated action. When this is the case, EGAS generates additional candidates from the current policy and selects the action with the highest estimated probability of success among the initial action and candidates with sufficient historical support. Only the selected action is executed. EGAS then uses the experience pool to construct pairwise preferences between the selected action and unexecuted candidates. An auxiliary objective incorporates these preferences into the base RL objective, enabling the policy to learn from unexecuted candidates without requiring additional downstream trajectories. EGAS improves held-out task success over GiGPO by 1.7–7.9 percentage points across ALFWorld, WebShop, and Sokoban. It also reaches GiGPO's best-observed performance while using 7.3–25.0% fewer completed training trajectories. These results demonstrate that comparing and learning from more actions need not require executing more trajectories.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.