RoboCTR: Cross-Modal Trajectory Rewarding for World Model-Based Robotic Planning
Abstract
Action-conditioned world models require a planning score that evaluates predicted trajectories with respect to the task goal. We propose RoboCTR, a Cross-Modal Trajectory Reward (CTR) interface that scores a latent rollout against the available text, goal-image, and demonstration-video goals. Stage I learns a goal-conditioned per-step critic; Stage II freezes this critic and uses its gradients to shape the V-JEPA2 predictor; receding-horizon MPC then reuses the same critic to rank candidate action sequences. We further introduce RoboTrain, an eight-task simulation collection with keyframe-aligned subtask instructions and positive and negative trajectory segments. On seven RoboTwin2 tasks, RoboCTR achieves 84.3% macro success versus 79.7% for 0, a gain of 4.6 percentage points, with the largest gains on Pick Dual Bottles and Place Object in Basket. Additional ablations examine the training stages, reward interface, goal modalities, task decomposition, and guidance weight. We also present qualitative demonstrations on a physical manipulator.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.