RECALL: Recovery Experience Collection for Active Lifelong Learning in Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models are typically trained offline on large demonstration datasets and updated post-deployment through additional fine-tuning. A more capable system should instead improve continually, actively identifying states in which it should collect new demonstrations and without forgetting previously learned behaviors. We present RECALL, the first empirical study of active, lifelong learning in VLAs, investigating two coupled questions: where to collect new supervision and how to incorporate it into a pretrained policy. We evaluate a VLA's ability to retain prior behaviors while adapting to new data, comparing start-state demonstrations against policy-induced recovery states, uncertainty-guided versus non-uncertainty-based state selection, sparse versus dense recovery collection, and continual adaptation strategies including reduced learning rates, elastic weight consolidation, and replay. Key findings include: recovery-state supervision is often more useful than repeated start-state demonstrations, though this advantage does not require uncertainty-based selection; more recovery data is not always better; targeted recovery data alone can cause severe forgetting; and preserving prior behavior requires maintaining appropriate behavioral coverage. We further demonstrate uncertainty-triggered recovery collection as a pre-execution intervention on a physical robot. Together, these findings establish an empirical foundation for active lifelong learning in VLAs and offer practical guidance for building robot policies that improve continuously from deployment experience.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.