GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
Abstract
World-action models (WAMs) commonly adapt video generators pretrained on natural video, leaving open how their manipulation capability scales when they are pretrained from scratch on manipulation data. To study this question, we introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model that combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE learns compact representations for robot control, while the SVP and IDM are pretrained from scratch on manipulation data. Knowledge-aligned selective optimization (KASO) jointly trains the SVP and IDM using generated futures compatible with recorded actions. These components generate a 52-action chunk in 104 ms on a single RTX 5090, supporting real-time control at 30 Hz. We study zero-shot out-of-distribution (OOD) generalization without per-task fine-tuning on 100 real-robot tasks across 20 manipulation skill groups and two embodiments. Scaling co-training data from 300 to 30,000 hours substantially improves success rates on both G1-OP and G2-90D, despite the latter contributing less than 2% of the training mixture. Beyond total data scale, our skill-level coverage analysis reveals a strong correlation between skill-specific training hours and zero-shot OOD success. We further study fine-grained instruction following, achieving ≥90% Follow Score for object identity, color, shape, and position. On RoboTwin-Clean2Rand and GenieSim-Instruction, GE-Act 2.0 outperforms strong recent baselines including π0.5, further demonstrating its generalization ability. We will release pretrained checkpoints and code.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.