acceptodds
Under review as a conference paper at ICLR 2027

LinguaMotion: Expressing the Language of Motion for Real-World Robot Manipulation

Abstract

Translating high-level instructions into manipulations is a fundamental capability in robotics, yet achieving this across diverse tasks and embodiments in real world remains challenging. While Vision-Language Models (VLMs) offer rich world knowledge, existing VLM-based policies often couple action prediction to specific robot kinematics and action spaces, failing to fully leverage pretrained knowledge and limiting the generality. We introduce LinguaMotion, a novel framework reinterpreting manipulation as structured generation in the language of motion. Instead of low-level robot actions, our model represents manipulation via object-centric motion directly within the native language space of VLMs. Given an image and instruction, it generates a hierarchical description of task-relevant objects, their spatial states, and how they should evolve to accomplish the task, thus bridging high-level semantic understanding with physical motion planning. To support this formulation, we further introduce a numeric value loss to restore geometric fidelity in tokenized motion representation, and a cycle-consistent learning framework to enforce semantic alignment between motion and language while enabling inference-time self-verification. On challenging real-world manipulation tasks, LinguaMotion consistently outperforms mainstream generalizable baselines across embodiments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.