VLMs Play Robotics: Can VLMs Control Locomotion and Manipulation?
Abstract
Vision-language models (VLMs) can reason over images and write code, but whether these abilities transfer to robotic control is unclear. To answer this question, we define several interfaces—interaction patterns—that correspond to different resources that VLMs commonly have access to during deployment. We then evaluate frontier and open-weight models on those interfaces across toy tasks, quadrupeds, humanoids, and robotic arms. We find that models are stronger at higher-level control, although nearly every capability we test shows generational performance gains within each model family. Specific interface and tool choices change performance by an order of magnitude, meaning narrow evaluations may under- or overestimate a model’s ability to exert control. Though overall success rates are low and other factors such as latency prevent effective deployment, our experiments indicate that generalist VLMs possess nascent robotic capability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.