Aligning Vision-Language Models for Goal Planning in Reinforcement Learning
Abstract
Exploration remains a major challenge for reinforcement learning (RL) agents in Craftax-style environments due to sparse rewards and long horizons. LLM-guided RL is a promising direction, but existing methods still suffer from substantial computational and financial costs as well as training instability due to their reliance on frequent queries to large-scale LLMs or carefully designed task-specific prompts. In this paper, we focus on developing an alternative approach that offers strong performance, stability, and cost-effectiveness. Specifically, we construct CCGP-2M, a multimodal goal-preference dataset for the Craftax-Classic environment, and propose a VLM-guided RL framework that fine-tunes a lightweight VLM to guide the RL agent's exploration. This framework consists of an offline stage, in which a Kullback-Leibler Token-level Direct Preference Optimization (KLTDPO) method is introduced to align the VLM with the goal preferences in CCGP-2M, and an online stage, in which the fine-tuned VLM guides the RL agent's exploration. Experiments on Craftax-Classic show that the proposed approach outperforms state-of-the-art baselines in terms of both score and reward, while achieving the highest success rates on 15 out of 22 achievements and the second-highest success rates on another five.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.