Disentangling Task-Relevant Latent Dynamics for Reward Modeling
Abstract
Reward design remains a major bottleneck for reinforcement learning in robotic manipulation. Handcrafted dense rewards can compensate for the limited guidance provided by sparse rewards, but require substantial task-specific reward engineering. Learning rewards from videos provides a scalable, data-driven alternative to manual reward design. However, existing video-based rewards are often computed in entangled visual spaces where task-relevant state changes are confounded with extraneous visual variation. As a result, the learned reward may respond to visual distractors rather than actual task execution. To address this limitation, we propose DiTaR (Disentangling Task-Relevant Latent Dynamics for Reward Modeling), a framework for learning dense rewards from heterogeneous videos. DiTaR uses an asymmetric two-stage representation learning procedure to separate task-relevant state information from task-irrelevant visual variation. In the task-relevant latent space, we train a history- and language-conditioned rectified-flow model to capture task dynamics and derive a directional dense reward from temporal changes in dynamics compatibility. Together, these components enable reward modeling directly over task-relevant dynamics rather than generic visual variation. DiTaR substantially improves policy learning on MetaWorld-v3 and ManiSkill over video-based reward baselines, and further supports real-world offline policy learning through reward relabeling. Reward analyses further show that DiTaR produces a more reliable reward signal for robotic policy learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.