acceptodds
Under review as a conference paper at ICLR 2027

Sim2Critic: Learning Reward Critics from Gaussian Splatting Simulation for Real-World Robot Manipulation

Abstract

Simulation offers a key advantage for robot learning: privileged access to the underlying state enables accurate supervision of task success and progress. On real robots, however, this state is hidden behind visual observations, leaving post-training of generalist policies dependent on sparse, human-labeled outcomes. We introduce Sim2Critic, a vision-based critic trained with privileged supervision in a 3D Gaussian Splatting (3DGS) digital twin of the deployment scene and transferred to real-world observations. We study two variants that differ only in their supervision: Sim2Critic (Sparse) learns from sparse success signals, whereas Sim2Critic (Dense) learns from dense progress signals derived from the same simulator state. Across eight real-world manipulation tasks on two robot platforms, both variants transfer to real rollouts, requiring as few as ten real demonstrations per task. It outperforms existing generalist reward models in reward alignment and outcome detection, while using substantially fewer parameters and achieving lower inference latency. Beyond evaluation, the transferred critic supports policy improvement without per-episode reward annotation. It defines returns for offline reinforcement learning on a policy’s own real-world rollouts and provides both rewards and success detection for online reinforcement learning on the robot. Through these two applications, we enhance the success rate of deployed policies by up to 2.2 times. Together, these results show that Sim2Critic turns a scanned workspace into a practical source of per-step reward for real-world robot learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.