HIVE: History-Initialized Visuomotor Evolution with a Compact Visual Bottleneck for Robotic Manipulation
Abstract
World-action models jointly predict future actions and visual outcomes, but noise-initialized generation can require costly iterative refinement. While some policies initialize generation from semantically informative priors, they predict actions alone without jointly forecasting visual outcomes. We present HIVE (History-Initialized Visuomotor Evolution), a flow-matching framework that uses historical actions and observed visual outcomes as the starting state for joint future prediction. This structured prior starts closer to the target than Gaussian noise, enabling more accurate generation with fewer integration steps while retaining the predictive visual supervision of world-action models. To keep joint modeling efficient, we introduce an Action-Centric Visual Bottleneck (ACVB), a lightweight 28.5M-parameter compressor that reduces 196 frozen DINOv3 patch tokens to very few compact, manipulation-informative tokens per frame. Dense-feature and policy supervision preserve local interaction information in a shared representation for history and future targets. Causal temporal attention captures action–outcome dependencies. Across 18 RoboTwin tasks, six-step HIVE improves mean success by 12.06 percentage points over its action-only prediction variant, and by 13.84 and 9.34 points over six- and ten-step noise-initialized variants, respectively. Compared with the ten-step variant, it reduces mean inference latency by 56.3% per eight-action chunk. Using ACVB instead of global CLS features improves mean success by 17.89 points over HIVE-CLS. It achieves 71% mean success on five real-world tasks, supporting effective and efficient history-initialized visuomotor prediction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.