ActEffect: From Predictive Foresight to Consequence Feedback in Robot Learning
Abstract
Robot actions are interventions on the physical world, and their quality is ultimately determined by the state transitions and task outcomes they induce. However, most robot policies are trained to match demonstrated actions, while predictive world models typically provide foresight for action generation rather than feedback on the outcomes of generated actions. We present ActEffect, which augments policy learning with a world model that predicts the consequences of proposed actions and feeds them back to guide policy optimization, internalizing awareness of action outcomes into the policy itself. To make this feedback effective for policy refinement, we develop a decoupled consequence modeling framework that predicts future transitions in a separate DINOv3 latent space, preserving spatial and structural information while reducing interference from task-specific semantics. We then introduce a comparative consequence refinement objective that uses the predicted consequences of coarse-to-fine action proposals from our policy to construct a relative feedback signal. By comparing the refined proposal against the best auxiliary consequence, ActEffect guides policy refinement toward actions whose predicted futures better align with the demonstrated outcome. The world model is used only during training and removed at deployment, introducing no inference overhead. Experiments on simulation benchmarks demonstrate that ActEffect achieves state-of-the-art performance while delivering a 2.5 inference speedup over the strongest baseline, reaching average success rates of 99.0% on LIBERO and 67.5% on RoboCasa-GR1. On real-world contact-rich manipulation tasks, ActEffect further outperforms the baselines by 13.5 percentage points in success rate while demonstrating strong robustness under perturbations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.