What If? Counterfactual Simulation for Sample-Efficient Visual Reward Learning
Abstract
Pretrained visual robot reward models capture generic notions of task success, not necessarily a person's specific preferences for how their task should be done. The robot can learn these preferences from user demonstrations, but with only a few demonstrations, a reward model may latch onto spurious correlations without learning what the person actually wants. To better express their preferences, a person could directly specify a preference axis—or an aspect of the task that they care about—in language. But a language-conditioned reward model may still reward the wrong behavior when training demonstrations don't fully cover the entire preference axis. Our key idea is to use the person's instruction to cover this axis by imagining what they don't want. We introduce COVER, which uses a video world model to generate counterfactual trajectories: alternatives that are not preferred along the preference axis but keep the rest of the behavior and scene the same. Comparing the original demonstration against its counterfactual isolates the quality that should make the reward different, without asking the person to demonstrate both versions. We train the reward to score each demonstration higher than its corresponding counterfactual, alongside learning from the demonstrations directly. Across four simulated and two real-robot tasks, COVER learns rewards that capture user preferences better than language-conditioned baselines under the same demonstration budgets, matching the strongest baseline's accuracy with up to 10× fewer demonstrations. On a real robot, selecting training demonstrations with these rewards yields policies that follow both stated preferences of each task more closely than training on unfiltered demonstrations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.