acceptodds
Under review as a conference paper at ICLR 2027

Action-Manifold VLA: Geometry-Aware Flow Matching in Joint and Task Spaces

Abstract

Diffusion- and flow-based action experts have become a dominant paradigm for vision-language-action (VLA) models, yet they hinge on two design choices: what geometry the action representation should obey, since a bimanual action chunk naturally lives on a product manifold rather than in Euclidean space, and in which space the generative dynamics should evolve, where joint space and task space each offer trade-offs. To study these questions, we propose Action-Manifold VLA (AM-VLA), a framework with two variants that share the same backbone and a unified action space but differ in where the flow lives. AM-VLA-J builds a joint-space flow on joint residual and gripper channels while introducing manifold-aware task-space auxiliary supervision, whereas AM-VLA-E builds a task-space flow on the per-arm end-effector residual manifold , modeling gripper, position residual, and rotation residual; rotations use shortest-geodesic conditional paths and Log-map target fields with time-constant norm. On the 50-task RoboTwin 2.0 benchmark, AM-VLA-J and AM-VLA-E achieve average success rates of and , respectively, compared with for the backbone, showing that manifold-aware task-space supervision improves joint-space generation, while manifold-native dynamics provide substantially larger gains when the generative process evolves directly in task space.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.