StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
Abstract
Robot learning continues to explore different pretrained models, approaches to action modeling, and sources of supervision. These choices require infrastructure that accommodates new methods while providing reusable training and evaluation workflows. We present StarVLA, an open-source infrastructure that gives the robot learning community accessible, reusable models, training tools, and benchmark resources. Model-level frameworks let researchers compose pretrained components, including vision-language models (VLMs), with action heads and modify their connections and learning objectives. The infrastructure provides 27 registered policy implementations spanning visuomotor, VLM-based, and world-model-based policies. Scalable, extensible dataloaders and configurable dataset mixtures integrate robot data, human-centered data, and multimodal data. Shared training tools support robot post-training, cross-embodiment training, VLM training, and multimodal co-training. The 13 simulation benchmark integrations cover 10 robot embodiments, spanning single-arm, bimanual, humanoid, and mobile manipulation, as well as navigation. Through these integrations, StarVLA makes baseline training recipes, evaluation entry points, and at least 41 policy checkpoint repositories available for research. A shared policy-serving path supports simulation and real-robot clients while keeping environment dependencies separate from model execution. Community research builds on these resources to extend policy architectures and conduct representation pretraining. StarVLA provides infrastructure for researchers to develop their own models, training procedures, and evaluations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.