acceptodds
Under review as a conference paper at ICLR 2027

VLMs are Native Action Models: Unleashing the manipulation capabilities of VLMs with Explicit Action-Space Grounding

Abstract

Pretrained vision-language models (VLMs) already exhibit strong visual, spatial, and reasoning capabilities relevant to robotic manipulation. A key challenge in turning these capabilities into generalist manipulation policies is how to expose these capabilities through an appropriate action interface. Vision-language-action (VLA) models provide a unified trainable policy, but typically introduce learned action tokens, action heads, or action experts whose semantics must be acquired from robot trajectories before the model can act. In contrast, modular foundation-model-based manipulation methods can exploit pretrained VLM capabilities for zero-shot manipulation, but rely on external planning and control scaffolds that are not end-to-end trainable as a unified continuous-action policy. We introduce EASG, an Explicit Action-Space Grounding framework that bridges these two paradigms. EASG makes the semantics of a continuous robot action space directly observable in the VLM’s perceptual context by projecting a tool center point (TCP)-centered action frame into first-person observations, while expressing actions through the VLM’s native textual generation interface. This allows pretrained VLMs to directly translate their existing visual-spatial reasoning into executable robot actions, without robot-action pretraining or auxiliary learned modules. At the same time, the same interface supports direct fine-tuning with robot trajectories, allowing the model to continuously improve upon the manipulation capabilities unleashed by EASG. We demonstrate that EASG enables zero-shot manipulation and generalization across object positions, categories, and robot embodiments, while remaining effectively trainable in continuous action spaces. These results suggest that, given an appropriate action interface, pretrained VLMs can themselves serve as native action models for robotic manipulation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.