GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
Abstract
Recently, Group Relative Policy Optimization (GRPO) has been widely adopted to train vision-language model (VLM) agents for diverse tasks. However, the context lengths of trajectory samples are too long for training VLM agents with GRPO to execute open-world tasks, which usually consumes too many training resources to remain practical. To address this issue, we propose GROW, a reinforcement learning (RL) framework for open-world VLM agents that decomposes trajectory samples into state-action samples, whose context lengths are shorter than those of trajectory samples, thereby reducing the required training resources. After decomposition, GROW performs credit assignment with temporal discounting on state-action samples, thereby converting sparse reward signals into dense ones to guide efficient learning, and it then computes advantages among the state-action samples within the same rollout groups. We further provide a surrogate analysis indicating that GROW retains the training signals for the whole trajectory samples in the optimization objective that can reinforce the successful trajectory samples and suppress the failed ones, while avoiding direct training on trajectory samples. Experiments on more than 800 Minecraft tasks show that our method can achieve higher average success rates and fewer steps than recent state-of-the-art approaches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.