WAVE: Learning to Verify Actions in 3D World Space for Generative Robot Policies
Abstract
Generative robot policies, including vision-language-action models (VLAs) and world-action models (WAMs), produce temporally extended action chunks from observations and instructions. At deployment, native sampling exposes multiple candidate chunks, yet standard control commits to one without comparing how the candidates would change the scene. We introduce WAVE (World-space Action VErification), an overall policy enhancement for generative robot policies. WAVE reconstructs the current scene once as a dense Gaussian field, compresses it into a compact representation, and predicts a separate compact future rollout for every candidate chunk. A unified score fuses the policy's native consistency score with learned verifier scores and structured verifier adjustments in the shared world-space frame. Across three simulation benchmarks and nine matched frozen-policy comparisons, WAVE reaches 99.0% on LIBERO, 79.5% on RoboCasa, and 95.6% average success on RoboTwin 2.0. These results show that consequence comparison converts native candidate diversity into reliable manipulation decisions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.