PatchVLA: Process-Aware Test-Time Critic-Guided Search for Long-Horizon Manipulation
Abstract
Reliable action selection is essential for completing long-horizon robotic manipulation tasks. However, comparing candidates from generative vision-language-action (VLA) policies remains challenging because plausible action chunks from the same observation can imply different interaction timelines. To address this challenge, we propose PatchVLA, a process-aware framework for critic-guided action search at test time. A process-conditioned flow policy generates candidate chunks, each paired with temporally aligned forecasts of interaction phase, local progress, and subtask boundaries. A sequence critic evaluates the actions and forecasts together with the current observation. Calibrated critic scores and auxiliary penalties guide MPPI-like updates to the latent proposal distribution while learned weights remain fixed. The controller executes a short prefix of a selected candidate and replans from the next observation. Process forecasts are learned from demonstrations with expert anchors and human-reviewed VLM annotations, while recorded candidate outcomes supervise the critic. On LeHome-Fold and LIBERO-Long, PatchVLA achieves 76.25% and 96.20% success, respectively, exceeding the official π0.5 baselines by 9.17 and 2.80 percentage points. In a separate fixed-model LeHome-Fold study, two-round refinement achieves 73.96% success versus 70.21% for one-round selection. Together, these results show that candidate-aligned process forecasts and critic-guided latent search improve long-horizon action selection. The anonymous project page is available at https://patchvla-anonymous.github.io/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.