Beyond Successful Rollouts: Learning from Simulator Near-Failures
Abstract
Simulator-verified near-failures expose physical distinctions that successful trajectories leave implicit. We show that supervising these distinctions improves manipulation reasoning beyond the benefit of successful-rollout exposure. Contrastive Embodied Reasoning (CER) rewards intermediate reasoning representations for distinguishing task-matched successful states from plausible physical failures. On CRD-Bench, a 30-task RLBench-derived diagnostic, CER exceeds outcome-only training by 10.9 points on high-complexity tasks after successful-trajectory exposure is matched (95% CI [8.3, 13.1]). Comparisons using the same simulator evidence establish the benefit of process supervision; an independent factorial isolates a further 2.53-point contribution from contrastive normalization (95% CI [0.5, 4.6]). The learned constraint representations affect decisions: removing their subspace lowers accuracy by 8.5 points, versus 1.1 for a rank-matched random subspace. The advantage over strengthened terminal training persists on 20 additional high-complexity classes (+9.6 points), independently constructed distractors, and unseen layouts. In executable RLBench planning, CER raises success from 25.3% to 35.3%, with comparable gains from non-contrastive process supervision. Verified near-failures provide a practical route from outcome verification to process supervision for manipulation reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.