FORGE: Feedback-Optimized Runtime Guidance Engine for Agentic RL
Abstract
Reinforcement learning (RL) improves language agents through interaction, but sparse terminal rewards can leave difficult tasks without informative training feed-back. In Group Relative Policy Optimization (GRPO), an all-failure rollout group yields zero task-reward advantages, removing its contribution to the task-directed policy gradient. Guidance can open new paths to success, yet its value for train- ing depends on whether the resulting experience improves the students subsequent unassisted behavior. We introduce FORGE( Feedback-Optimized Runtime Guidance Engine) that trains a harnessed guidance agent alongside the student using post-update learning value. When a clean rollout group entirely fails, FORGE launches candidate tutoring interactions, adapting its guidance to the students evolving execution history. An episode workspace and a retrievable experience library support decisions about what guidance to provide, when to intervene, and when to end assistance. To evaluate each interaction, FORGE compares paired temporary student updates with and without its trajectory group, measuring the resulting loss difference on independent, unassisted probe trajectories. This esti- mates the candidates contribution to student learning beyond the common training data. The resulting feedback trains the guidance policy over complete interactions, including retrieval and follow-up decisions. All eligible candidates contribute to guidance-policy training, while at most the highest-value positive group enters the students task-reward update. Experiments on multiple interactive agent benchmarks show that FORGE outperforms GRPO and prior guidance-based baselinesin unassisted task performance, demonstrating the effectiveness of training dy-namic guidance with post-update learning value.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.