Beyond Action Diversity: Effect-Class Exploration for LLM Agents
Abstract
Reinforcement learning can improve language model agents on interactive tasks, but limited training budgets are often spent continuing rollouts from different actions that produce the same immediate environmental outcome. We aim to convert repeated environmental outcomes into more effective exploration without treating distinct language histories as interchangeable. We introduce Effect-Class Exploration (ECE), which probes candidate actions in deterministic environments with restorable states, groups identical complete one-step outcomes, and continues one uniformly sampled representative per group. Each representative keeps its own history and receives a weight equal to its group size, preserving the expected additive policy gradient; our fixed-policy analysis identifies when interaction savings outweigh the additional sampling variance. Under matched interaction budgets, a fixed node bank, and a shared training update, ECE improves Qwen2.5-7B-Instruct success rates over development-selected direct sampling by 3.2 percentage points on unseen ALFWorld tasks and 5.1 points on WebShop, and exceeds optimized parsed-action grouping by 1.9 percentage points in WebShop success rate. ECE allocates exploration by observed effects under explicit efficiency conditions, with ordinary single-trajectory evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.