PONIX: Pseudo-Task-Free Continual Adaptation of Vision-Language-Action Models
Abstract
Vision-language-action (VLA) models must continually adapt to new tasks while preserving performance on previously learned tasks. Task-directed robotic data collection groups demonstrations of the same intended task, but assigning globally consistent task identities across collections requires additional annotation and coordination. We formulate *pseudo-task-free continual learning*, where demonstrations arrive in task-homogeneous mini-batches without oracle task identities or task boundary signals. We introduce **PONIX** (Pseudo-task-free ONline Inferred eXperts), an exemplar-free method that exploits this local structure. **PONIX** uses frozen multimodal representations and online routing statistics to select among action-side low-rank adaptation experts, exploiting within-batch task homogeneity to guide expert expansion. Evaluations across LIBERO and Meta-World with GR00T, , and SmolVLA show that **PONIX** approaches oracle-task-ID performance in both average acquisition success and forgetting mitigation. On revisitation experiments, **PONIX** achieves comparable average acquisition success to adaptation using task boundaries while maintaining approximately half as many experts. Real-world robot experiments further show that **PONIX** distinguishes different tasks expressed by the same instruction and reuses an expert for the same task expressed using different instructions. These results demonstrate continual VLA adaptation through inferred expert allocation and reuse, without oracle task identities, task-boundary signals, or stored exemplars.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.