TRACK: Object Keypoints and End-Effector Waypoints as Geometric Task Representations
Abstract
Executing a new manipulation task requires translating a task specification into actions suited to the robot. Geometric specifications obtained from demonstrations or generated by planners can describe the intended behavior, but their hand or end-effector trajectories generally cannot be replayed directly to accomplish the task. We propose TRACK, a framework for specifying and executing manipulation tasks through object keypoint trajectories, abstract grasp flags, and approximate hand or end-effector waypoints. Object trajectories and grasp flags specify desired object motion and grasping interactions, while hand waypoints provide soft guidance that allows deviations needed for successful manipulation. This shared representation accommodates specifications from scripted experts, motion planners, and vision-language models. We train a policy across multiple tasks using reinforcement learning, with reference-derived rewards for object motion, grasping, and end-effector tracking within a positional tolerance, without task-specific success rewards. A state-based progress tracker aligns the policy’s reference window and training reward with the robot’s execution progress. At test time, the policy receives a single task specification and executes the task without additional training. Experiments on single-arm and bimanual manipulation demonstrate generalization to unseen tasks from a single task specification. TRACK outperforms reference replay with approximate end-effector trajectories and imitation learning-based baselines, and also executes specifications generated by a vision-language model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.