Making Every Interaction Count: Efficient Real-World RL for Robot Policies
Abstract
Real-world reinforcement learning (RL) is crucial for adapting robotic control policies to physical environments. Yet, it faces several challenges: costly environment interactions and human supervision, sparse rewards in long-horizon tasks, and underutilized suboptimal experience from autonomous policy rollouts. To address this, we introduce **D**etect, **E**valuate, and **C**ompose (DEC), an RL framework that minimizes human intervention while enabling efficient autonomous policy improvement. DEC first trains a dynamics-aware discriminator to detect failures and determine when to trigger human takeover and resume autonomous execution. It then evaluates a robust value function to tackle the sparse reward challenge by adopting trajectory-held-out bootstrapping and discounted partial returns. For policy optimization, DEC learns complementary positive and negative sub-policies to exploit both successful and suboptimal experience, and compose them at inference time to enable controllable action generation. Across simulated manipulation environment and long-horizon, dexterous bimanual tasks in the real world, DEC consistently outperforms existing RL and policy fine-tuning methods with limited human intervention, demonstrating that effective real-world RL can make every interaction count. More results can be seen at https://iclr27-4081.github.io/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.