RoboCritic: A Grounded Reward Model to Assess and Guide Robot Policies
Abstract
Reliable feedback is critical for reinforcement learning-based robot policy improvement and deployment-time guidance. Terminal rewards summarize completed rollouts, while progress estimates based on temporal cues can diverge from the observed state during stalls or regressions. Scalar scores also leave the underlying physical conditions and appropriate next steps implicit. We introduce RoboCritic, a vision-language reward model that can be queried as observations arrive to assess execution and provide grounded guidance. Given a task instruction and one or more camera streams, RoboCritic maintains bounded causal memory with persistent task context and the initial observation, a sliding visual window, and a history of generated feedback. This design retains execution context without storing the full observation history. At each query, it jointly outputs a progress score, a grounded state description, and actionable next-step or corrective feedback. The final query recovers terminal outcome assessment. We train RoboCritic on RPF-1M, extending RoboReward trajectories with over one million aligned process-feedback records for joint supervision. The frozen critic supplies progress-difference rewards for reinforcement learning and language instructions to a frozen vision-language-action policy. Adding state descriptions and guidance as both supervision and history reduces progress-estimation MAE from 0.441 to 0.404. On two real-world manipulation tasks, RoboCritic achieves higher success than the compared baselines in reward-guided policy learning and language guidance of a shared frozen policy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.