Online VLA Adaptation with Streaming Reinforcement Learning
Abstract
Streaming Reinforcement Learning (RL) learns directly from sequential experience without relying on replay buffers or minibatch updates. It offers an appealing framework for online adaptation of vision-language-action (VLA) models, but this direction remains largely underexplored. In this work, we investigate this problem and develop Streaming with Intentional and Episodic Updates (SIE), a new streaming VLA adaptation framework that combines streaming critic learning with an Episodic Policy Updates (EPU) mechanism, which incorporates episodic success feedback through a lightweight actor update. Experiments with on LIBERO-90 demonstrate that our framework improves mean final success rate from 47.20% to 72.09% compared with the baseline. In a synchronous two-device benchmark, SIE reduces client wall time by 66.5%. Our results show that streaming RL is a promising direction for efficient VLA adaptation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.