Reward as an Agent for Embodied World Models
Abstract
RL for world models largely relies on conservative rollouts near the training distribution, limiting exploration and behavioral diversity. In this work, we challenge this conservative paradigm. We argue that the core limitation is not exploration itself, but the lack of reliable verification strategies to support broader exploration. Without reliable verification, policies can exploit proxy rewards through convincing videos that lack evidence of genuine task progress. On the verification side, we introduce Reward as an Agent, which centers evaluation on evidence-grounded task verification. It separates observed behavior from fixed task requirements, actively inspects disputed evidence, and revises judgments through tool-assisted reflection and requirement checks. Reward gates preserve supported progress while withholding credit for failure or unverifiable essential content. The framework is extensible: additional reward strategies can be integrated as tools to enrich verification and support further performance gains. On the exploration side, we introduce Dynamic-Aware Rollout Diversification through DynDiff-GRPO, which allocates stochastic exploration to dynamically salient regions to diversify trajectories. Together, these designs couple broader exploration with evidence-based reward evaluation. Training experiments demonstrate benchmark improvements across three open-source world models, while complementary offline evaluations show that Reward as an Agent improves reward reasonableness and reduces erroneous positive rewards on failure and unobservable clips.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.