acceptodds
Under review as a conference paper at ICLR 2027

SpecRL: Specification-Guided Dense Rewards for VLA Reinforcement Learning

Abstract

Reinforcement learning fine-tuning (RLFT) has emerged as a promising approach for adapting pretrained vision-language-action models through task-directed interaction. However, many existing methods rely on sparse binary rewards, which provide limited guidance for policy optimization and thereby hinder learning performance. While recent approaches seek denser feedback via LLM-generated reward functions or vision-language model evaluations, they are often inefficient and the resulting rewards are often unreliable, potentially introducing errors that misguide policy optimization. In this work, we propose SpecRL, a specification-guided framework designed to provide dense, interpretable online rewards for RLFT. The core idea is to formalize task requirements in Signal Temporal Logic (STL) and derive rewards from its explicit quantitative semantics. SpecRL first uses an LLM-aided offline process to convert language instructions into STL specifications constituted by task-level atomic propositions. During execution, SpecRL evaluates proposition satisfaction using either a privileged-state grounding module or a lightweight visual grounding module, supporting both simulation and real-world deployments. A runtime monitor then aggregates the quantitative values of propositions along the specification's temporal structure, translating fine-grained task progress into dense rewards. Extensive evaluations on the LIBERO benchmark show that privileged-state SpecRL improves the success rate over binary-reward RLFT by percentage points, while visual SpecRL surpasses the strongest visual baseline by percentage points with faster training. Real-world trajectory evaluations further demonstrate that its deterministic and interpretable design delivers more accurate reward feedback than baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.