CRL-VLA: Continual Vision–Language–Action Learning
Abstract
Lifelong learning is critical for embodied agents in open-world environments, where reinforcement learning fine-tuning has emerged as an important paradigm to enable Vision-Language-Action (VLA) models to master dexterous manipulation through environmental interaction. Thus, Continual Reinforcement Learning (CRL) is a promising pathway for deploying VLAs in lifelong robotic scenarios, yet balancing stability (retaining old skills) and plasticity (learning new ones) remains a formidable challenge. We introduce CRL-VLA, a framework for continual post-training of VLA models with theoretical performance-change bounds. The bounds relate old- and new-task return changes to goal-conditioned advantage magnitude and policy divergence. Motivated by this sensitivity analysis, the revised CRL-VLA formulation combines a frozen retention action critic, a trainable adaptation action critic, and policy regularization for retention and adaptation. Evaluations in simulation and on physical robots show that CRL-VLA supports continual adaptation across diverse VLA RL post-training paradigms and robot embodiments while maintaining robust task performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.