VLA-Collect: Training Vision Language Action Models with Few-shot Data via Collective Action Fine-Tuning
Abstract
Vision-Language-Action (VLA) models unify perception, language, and action and achieve strong performance across robotic tasks, but high-fidelity dexterous manipulation often still relies on costly supervised fine-tuning (SFT), pushing deployment into a few-shot regime. With few demonstrations, a policy can overfit to narrow behaviors and encounter unfamiliar states during execution. When self-generated rollouts are reused for training, low-quality or repetitive trajectories can reinforce execution errors and limited coverage. We propose Collective Action Fine-Tuning (CoAct), a collective-learning paradigm for interactive few-shot VLA post-training using limited initial expert demonstrations, additional rollouts, and binary task-completion feedback. CoAct maintains a diverse policy pool with two mechanisms: WeightCoAct encourages action-head discrepancy, while DataCoAct assigns persistent expert-specific visual views to broaden the experience generated by the pool. CoAct further introduces Collective Trajectory Curation (CoTC) to select successful rollouts and favor experience beyond the initial demonstration coverage, then distills curated trajectories into a unified student policy. Experiments on LIBERO, LIBERO-Plus, and RoboTwin, complemented by five real-world tasks, show improved few-shot performance and zero-shot transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.