acceptodds
Under review as a conference paper at ICLR 2027

MARVL: Multi-Stage Guidance for Reinforcement Learning via Vision-Language Models

Abstract

Designing dense reward functions is pivotal for efficient Reinforcement Learning (RL). However, most dense rewards rely on manual engineering, which fundamentally limits the scalability and automation of reinforcement learning. While Vision-Language Models (VLMs) offer a promising path to reward design, naïve VLM rewards often misalign with task progress, struggle with spatial grounding, and show limited understanding of task semantics. To address these issues, we propose **MARVL**—**M**ulti-st**A**ge guidance for **R**einforcement learning via **V**ision-**L**anguage models. MARVL fine-tunes a VLM for spatial and semantic consistency and decomposes tasks into multi-stage subtasks with task direction projection for trajectory sensitivity. Empirically, MARVL significantly outperforms existing VLM-reward methods on the Meta-World benchmark, demonstrating superior sample efficiency and robustness on sparse-reward manipulation tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.