acceptodds
Under review as a conference paper at ICLR 2027

Layered Value Forecasts for Online RL Fine-Tuning of Vision-Language-Action Models

Abstract

Vision-language-action (VLA) models provide strong priors for robot manipulation, but adapting them to a new task with online reinforcement learning (RL) is costly, since every episode has to be run on a real robot. Part of this cost comes from the reward, which is sparse and arrives only when the task succeeds. A decision such as closing the gripper on an object is credited only through the value of the states that follow it, and that value stays near zero until the policy completes the task from them, so the critic separates a good decision from a bad one only once success is already frequent. Intermediate events, in contrast, occur in failed episodes as well as successful ones. To exploit this, we propose Layered Value Forecasts (LVF), which augments the critic of an off-policy actor-critic with event-based general value functions (GVFs). Each GVF predicts the discounted probability that one event occurs, such as a grasp or a collision, and is learned from binary event labels with minimal supervision. Their prediction at the chosen action enters the critic's TD target as a cumulant alongside the terminal reward, so the critic receives credit at the steps leading to each event. Across three real-robot manipulation tasks, LVF improves most where success is rare, reaching 95.0% and 70.0% on the two tasks where the initial policy succeeds in 32.5% of trials.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.