TriBrain: A Three-System Embodied Foundation Model for Scalable Robot Manipulation
Abstract
Vision-language-action (VLA) models have become a leading approach to generalist robot control, yet unifying understanding, prediction, and control within an embodied foundation model remains challenging. In particular, staged VLA training separates vision-language adaptation from continuous-action learning, limiting the co-adaptation of semantic representations and continuous control. Moreover, integrating world models into VLA pipelines for data augmentation or visual subgoal prediction, without coupling future-state prediction with task-progress evaluation, limits their effectiveness in guiding long-horizon manipulation. To address these issues, we present TriBrain, a multi-embodiment VLA pretraining framework. The model adopts a three-system architecture: System 1 converts multimodal context into continuous robot actions; System 2 performs scene understanding and task decomposition; and System 3 predicts future observations and estimates task progress to guide action generation. TriBrain jointly trains a pretrained VLM backbone and a continuous-action expert while incorporating multimodal understanding objectives, mitigating optimization fragmentation. Soft Knowledge Insulation attenuates action gradients entering the vision-language backbone to balance semantic preservation and motor adaptation. During policy post-training, a frozen World Value Model (System 3) provides predicted subgoal images and task-progress estimates used to derive binary advantage labels. The subgoal images and advantage labels serve as policy inputs and are independently dropped during training. Trained on over 37,000 hours of heterogeneous embodied experience, TriBrain surpasses in real-world manipulation, achieving 84.9% and 74.1% success on complex tasks with AgileX PiPER and Maker H01, respectively. In simulation, it outperforms by 9.0% and 5.22% on RoboTwin 2.0 and EBench, respectively. In embodied vision-language evaluation on MiMo-Embodied, it also achieves state-of-the-art performance with an overall score of 0.4621. Code and pretrained weights will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.