SimVLA 2.0: Instruction-Conditioned Control without Language Pretraining
Abstract
Vision-Language-Action policies typically retain pretrained language components, yet on finite, recurring task inventories the role of that retention is rarely tested directly. Action training itself can erode the competence a pretrained backbone contributes, and what remains often fails to transfer under meaning-preserving paraphrases. We therefore build the opposite end of the design space: instructions remain, but no pretrained language-side component is loaded. SimVLA 2.0 loads no language model, text encoder, or token embedding: a fixed, parameter-free mapping indexes randomly initialized embeddings learned solely from action supervision, and a joint action Transformer conditions a continuous action head on the embeddings and the pretrained vision encoder's features. Across seven benchmark evaluations spanning six manipulation domains, it remains within the range of recent language-pretrained systems—LIBERO 96.9%, LIBERO-PRO 96.5% under semantic instruction perturbations, CALVIN ABCD a 4.24 mean chain length, WidowX 91.7%, Google Robot 82.0%, and RoboTwin 2.0 Easy/Hard 94.0%/87.2%. On the 2025 BEHAVIOR Challenge Public Validation leaderboard it places second among Standard-track entries. Ablations separate the sources of performance: instruction conditioning is essential, pretrained token embeddings provide no consistent advantage over action-trained codes, and strong visual pretraining remains important.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.