Learning Quality-Controllable Robot Policies with Iterative Hindsight Relabeling
Abstract
Robot policy evaluation often relies on a binary success signal for task completion. In real-world deployment, however, desired execution behaviors are also defined by quality attributes such as execution speed and motion smoothness, which need to be dynamically controlled according to task requirements. To address this, we introduce Quality-Controllable Policy (QCP), a robot policy that enables inference-time control over both individual quality attributes and their combinations. To extend the controllable quality range beyond initial demonstrations, we introduce an iterative hindsight relabeling loop: QCP collects online rollouts under target quality labels, adds successful trajectories to the dataset, relabels the accumulated trajectories in hindsight according to their measured quality, and fine-tunes the policy. Experiments on robomimic and RoboEval demonstrate that QCP consistently improves both task execution quality and success rates across iterations, producing successful rollouts that exceed the quality of the initial demonstrations. Finally, we validate QCP on real-world manipulation tasks, showing that controlling speed increases throughput in conveyor operations and controlling transport stability improves success when carrying wine glasses, showing that quality-controllable policies yield practical gains in physical deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.