acceptodds
Under review as a conference paper at ICLR 2027

PointAction: 3D Points as Universal Action Representations for Robot Control

Abstract

World Action Models (WAMs) leverage the broad visual dynamics captured by pretrained video diffusion models, offering a promising path toward generalizable robot manipulation. However, RGB-only rollouts leave 3D motion and fine-grained spatial constraints implicit, making action grounding ambiguous. Meanwhile, scaling action supervision across tasks and embodiments remains costly. We present PointAction, a framework that uses dynamic 3D points as an action interface between video prediction and robot control. PointAction adapts a foundation video model to jointly predict future RGB frames and dynamic 3D pointmaps. This interface expresses proposed robot motion in a common geometric form, which an embodiment-specific diffusion decoder translates into controls. Geometric supervision from videos enables the rollout model to learn beyond paired action data, while target-robot video adaptation and decoder training support transfer with limited demonstrations. Experiments demonstrate strong 4D generation quality on robot scenes, competitive manipulation performance in simulation, and adaptation to two real robot arms unseen during pretraining. These results support 3D point dynamics as a common representation for connecting video prediction to robot control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.