acceptodds
Under review as a conference paper at ICLR 2027

Scaffolding General Vision Language Models for Zero-Shot Robot Control

Abstract

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization remains limited, and their reliance on specialized robotics data keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). Agentic robotic systems instead use VLMs for high-level reasoning, yet typically depend on substantial external modules. This motivates our question: Can a general-purpose VLM itself serve as the decision-making core for robotic manipulation? Evaluating VLMs on action selection, progress assessment, and task completion reveals a capability-control gap: VLMs reason well about manipulation, but their decisions do not match the abstraction of continuous, embodiment-specific control. We introduce MotorMind, a VLM-centric manipulation harness with an asynchronous inference loop. Without task-specific policy training, MotorMind achieves 66.7% base success and 53.8% success under perturbations on LIBERO-PRO, compared with at most 13.3% and 19.2% for the prior zero-shot methods we evaluate, and the same interface reaches 95% average success on a real xArm6 robot across direct manipulation and human-perturbation settings. A stronger VLM backbone further raises base success without changing the harness, and the remaining errors, mainly in visual grounding, embodied reasoning, and action knowledge, decrease with stronger VLMs. These results reframe zero-shot manipulation as aligning general-purpose model capabilities with the control representation, rather than solely learning specialized action policies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.