MARVL: Multi-Stage Guidance for Reinforcement Learning via Vision-Language Models
Abstract
Designing dense reward functions is pivotal for efficient Reinforcement Learning (RL). However, most dense rewards rely on manual engineering, which fundamentally limits the scalability and automation of reinforcement learning. While Vision-Language Models (VLMs) offer a promising path to reward design, naïve VLM rewards often misalign with task progress, struggle with spatial grounding, and show limited understanding of task semantics. To address these issues, we propose **MARVL**—**M**ulti-st**A**ge guidance for **R**einforcement learning via **V**ision-**L**anguage models. MARVL fine-tunes a VLM for spatial and semantic consistency and decomposes tasks into multi-stage subtasks with task direction projection for trajectory sensitivity. Empirically, MARVL significantly outperforms existing VLM-reward methods on the Meta-World benchmark, demonstrating superior sample efficiency and robustness on sparse-reward manipulation tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.