Can Cheap Failures Teach Good Rewards? Studying Counterfactual Sources for Reward Learning in Manipulation
Abstract
Reinforcement learning (RL) can improve robot policies beyond imitation learning, but it requires rewards that capture task progress and failure. Learning such rewards requires not only successful demonstrations, but also informative counterfactual failures that can guide the model on behaviors to penalize. However, large-scale robot datasets are dominated by successful behavior, and collecting failures on real hardware is expensive. This raises a practical question: given a limited data-collection budget, which sources of counterfactual experience are worth investing in? We compare four sources of failure data: teleoperated failures, policy rollouts, human videos, and real-to-sim rollouts, across five real-world manipulation tasks. Holding the task suite, successful demonstrations, and evaluation fixed, we study these sources through LoRA adaptation of a reward model and few-shot prompting of a frozen model, using matched-data and scaling settings. We evaluate the resulting reward models on a custom held-out benchmark of 250 human-annotated policy rollouts, measuring trajectory ranking, partial-completion calibration, and success detection. Finally, we use selected reward models to guide RL post-training of and test whether improvements in offline reward quality translate into better robot behavior. Together, our study examines whether inexpensive counterfactual sources can match or surpass smaller amounts of faithful real-robot failure data when scaled.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.