acceptodds
Under review as a conference paper at ICLR 2027

PrimingVLA: Self-Supervised Action Priming for Parameter-Efficient Vision-Language-Action Adaptation

Abstract

Adapting Vision–Language–Action (VLA) models requires learning task-specific control through an action expert while preserving the multimodal capabilities of a pretrained vision–language model (VLM). Existing methods primarily align multimodal representations with action generation, but overlook the need to establish structured action priors before task-conditioned learning. To address this issue, we introduce PrimingVLA, a parameter-efficient framework built on our proposed new principle of action priming before multimodal conditioning. Its Self-Supervised Action Priming (SAP) learns temporal action structure through masked trajectory reconstruction, after which multimodal inputs are introduced for joint adaptation. A cross-prompt calibration coordinates their adaptation by bounding relative action-prompt updates using VLM-prompt updates as a reference. Experiments on LIBERO, LIBERO-PLUS, and LIBERO-PRO demonstrate effective task adaptation and characterize robustness under distribution shifts, with ablations supporting the benefits of staged learning and coordinated updates. Real-world experiments on a Piper robotic arm further demonstrate its effectiveness in robotic manipulation. Neural Tangent Kernel monitoring throughout training provides direct evidence that the sensitivity gap between the two pathways progressively narrows. Remarkably, PrimingVLA requires updating fewer than 0.05M parameters ( of the full model) while keeping both pretrained backbones frozen, enabling efficient adaptation and rapid task switching.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.