PAVE: Separates Prediction from Preference in World-Action Policies
Abstract
An unsuccessful robot rollout records valid transitions alongside behavior a policy should improve. PAVE (Predictive Alignment and Value-Guided Evolution) equips world-action policies with trajectory-adaptive predictive alignment and action-conditioned value learning. Local and trajectory-relative joint-embedding targets train shared actor representations across temporal scales, retaining valid transitions from both successes and failures. A dual-head critic couples future-latent prediction with return estimation; episode-excluded values assign relative-quality conditions to action chunks. Both quality conditions train the actor, and deployment selects positive. Matched-data actor comparisons improve on local prediction with quality conditioning by 5.0/4.5 percentage points on LIBERO-Plus/RoboTwin Random. A matched critic ablation isolates the benefit of future-latent learning. On S1, PAVE achieves 85.7 ± 3.5%/92.0 ± 2.0% on organization/pouring versus 80.0 ± 1.0%/84.7 ± 1.5% for the matched control (mean ± sample SD over three training seeds; 100 trials per seed and task). Target encoding, predictive heads, and critics remain offline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.