Surpassing VLA Performance: Online RL Training of Lightweight Policies with Minimal Human Interventions
Abstract
Pre-trained vision-language-action (VLA) models often exhibit low success rates and task completion speeds when deployed without fine-tuning in new environments. The standard approach to improve the performance of a pre-trained VLA is to collect new demonstrations and fine-tune the VLA for every new task encountered during deployment. However, collecting human demonstrations during deployment can be costly and time-consuming, which is a barrier to the widespread adoption of VLAs in industry and daily life. In this work, we propose a framework based on reinforcement learning for surpassing the performance of pre-trained VLAs by learning a lightweight policy (LWP) with no human demonstrations and minimal human intervention. By reinforcement learning only from online experience, our proposed framework RELAY transfers task-specific robot behavior from the VLA to the LWP and improves the behavior over time. Our experiments on three real-robot manipulation tasks show that RELAY-based agents autonomously improve their performance and achieve higher success rates and task completion speeds than the pre-trained VLA within six hours of interaction time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.