ThoughtVLA: Evolving Thought Trajectories for Continual Vision-Language-Action Adaptation
Abstract
Continual adaptation of pretrained vision-language-action (VLA) models aims to acquire new manipulation skills from sequentially arriving demonstrations while retaining performance on previously learned tasks. However, in progressive LoRA-MoE methods, updating the router can change the expert mixtures used for previously learned tasks even when their experts remain frozen. Meanwhile, pretrained semantic similarity alone may not reliably distinguish learned skills that are semantically related but require different manipulations. To address these challenges, we propose ThoughtVLA, a continual adaptation framework built around Continually Evolving Thought Trajectory (CETT), which formalizes how pre-action reasoning for skill selection evolves as task knowledge accumulates. In ThoughtVLA, Evolving Skill Reasoning selects a learned skill by interpreting the current input with pretrained semantics and accumulated task knowledge, and resolving ambiguity with task attributes. Each acquired skill retains a LoRA composition, which is used throughout execution when that skill is selected. Beyond skill reasoning, Complementary Knowledge Transfer supports new-skill acquisition by selecting prior LoRA updates that reduce the residual error left beyond the limited rank of new adapter and incorporating them into the new skill’s LoRA composition. Across single-arm, bimanual, humanoid, mixed-embodiment, cross-backbone, and real-robot evaluations, ThoughtVLA achieves higher final success and lower forgetting than the compared methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.