Efficient-RLT: Sample-efficient Real-robot RL via Optimizing Human-intervention Strategy with Dense Process Rewards
Abstract
Real-world online reinforcement learning (RL) is a promising approach for post-training vision-language-action policies, particularly for precision-sensitive, long-horizon manipulation where supervised fine-tuning alone rarely achieves reliable execution. Existing human-in-the-loop methods intervene only after the policy has deviated, so the critical decision state never receives both a policy action and a corrective one. The critic therefore either extrapolates over unexecuted proposals or propagates the return of a successful recovery back through the very actions that caused the failure. We introduce Efficient RLT, a framework built on RL Token (RLT) that learns through action contrast from *rewind interventions*. When an operator flags an erroneous action chunk, the robot is physically rewound to the pre-error decision state by reverse-replaying its waypoints, and the operator issues a corrective action from there, yielding a state-aligned pair of a rejected action and its corrective replacement. Efficient RLT trains the critic and actor with a *pairwise preference learning* on these pairs, shaping a local gradient field over the action space that flows from rejected toward corrective actions, while branch-consistent credit assignment keeps the corrective return out of the rejected branch. We further train a progress estimator that supplies *dense progress rewards* under sparse task feedback, and study how takeover timing and horizon affect convergence. On Franka, Cobot, and X-Square platforms across diverse and challenging manipulation tasks, Efficient RLT improves sample efficiency, convergence speed, and final success rate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.