Crucible: Counterfactual RLVR For Dexterous Manipulation Planners
Abstract
Combining vision-language planners with control primitives is a promising paradigm for robotic manipulation. However, reliable execution still requires high-level planners to improve from the outcomes of their own actions, particularly when hallucinations or incorrect reasoning lead to failures. Learning from such feedback is challenging because a primitive may appear locally correct but cause failure many steps later, while a task-level failure signal does not reveal which decision should be corrected. This lack of credit assignment makes planner learning inefficient. We present Crucible, a reinforcement learning from verifiable rewards (RLVR) framework with counterfactual credit assignment. Crucible builds on a new vision-language-primitive (VLP) model constructed from robot demonstrations, with each primitive paired with a privileged verifier. For a failed trajectory, our counterfactual RLVR algorithm intervenes on individual primitive decisions from saved states and replays the recorded suffix, identifying decisions whose replacement repairs the observed verifier failure under fixed-suffix replay. Full-episode returns from these alternative executions then define advantages for masked PPO that updates only the identified decision tokens. We evaluate cross-step dependencies in tabletop grasping, where base motions affect the feasibility of a later grasp. A replay probe rescues 49.5% of eligible failures, demonstrating the availability of executable repairs. In closed-loop evaluation, counterfactual PPO achieves 62.38% overall success, compared with 55.45% for call-level episode-return PPO and 56.44% for the shared SFT baseline. Project Page: https://cruciblex2026.github.io/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.