acceptodds
Under review as a conference paper at ICLR 2027

RoboStep: Robot Rewards Benefit from Tracking Task Progress Step by Step

Abstract

Progress reward models aim to infer task progress from visual observations, especially when privileged task states are less accessible in the real world than in simulation. Yet progress is inherently task-specific and history-dependent, making a globally normalized score difficult to calibrate. Scalable supervision therefore often relies on frame positions in successful demonstrations, implicitly coupling progress with elapsed time and providing weak supervision for stagnation, regression, and recovery. In this work, we present RoboStep, a framework for constructing online robot rewards that explicitly tracks task progress step by step, using pretrained VLM semantics for step completion verification rather than global progress calibration. We use a Frame-VLM to decompose tasks and generate step-local dense reward functions, while a Gate-VLM decides when the reward advances to the next step and switches the active reward function. Across extensive simulated manipulation experiments, we systematically compare these two roles of VLMs under matched policy-learning settings. Combined with recent pipelines that reconstruct real-world scenes in simulation, RoboStep further enables policies trained with structured rewards in simulation to transfer back to corresponding real scenes. Overall, we show that robot learning benefits from tracking task progress step by step: for online reward, VLMs are better used as step-completion verifiers than as global scalar progress predictors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.