PhysCanvas: A Vision-Tactile World Model for Robot Policy Evaluation and Learning
Abstract
Evaluating and post-training robot policies through physical interaction is costly and difficult to scale, motivating action-conditioned world models as interactive simulators. However, visually plausible predictions can misrepresent contact dynamics and task outcomes, providing unreliable feedback for policy evaluation and learning. We introduce PhysCanvas, a state-consistent visuo-tactile world model for contact-rich robot manipulation. PhysCanvas organizes multi-view RGB, tactile images, marker displacement, and task state into a unified multimodal canvas. All modalities share a video VAE and pretrained generative backbone, allowing optical touch to reuse visual and temporal priors. Motion images derived through forward kinematics and camera projection provide spatially aligned action conditioning, while an initial RGB anchor and updated reference observations guide sequential generation. Our consistency-aware post-training uses the frozen Scene Reward Model, Modality Reward Model, and State Reward Model to supervise scene preservation, visual–tactile compatibility, and agreement with real task outcomes. The resulting world model supports closed-loop policy evaluation and provides imagined interactions and rewards for GRPO policy updates. Across four real-world manipulation tasks, State Reward Model supervision increases the arithmetic mean of task-wise Pearson correlations between model and real-world success rates from 0.602 to 0.885 under fixed-action replay. Policy post-training in the reward-guided world model further increases π0.5’s mean real-world success rate from a 59.1% SFT baseline to 70.8%. These results demonstrate the value of unified multimodal modeling and consistency-aware post-training for robot policy evaluation and learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.