RxBrain: Embodied Foundation Model with Joint Language-Visual Subgoal Planning
Abstract
We introduce RxBrain, an embodied foundation model that combines embodied understanding and reasoning with joint language-visual subgoal planning. It interleaves textual subgoals with goal images, pairing task decomposition and constraints with intended intermediate and final physical states. Built on vision-language model, RxBrain adopts a Mixture-of-Transformers architecture, with general image generation pretraining followed by interleaved text-image planning training. We introduce an inference-aware visual state construction strategy to enhance joint planning. Our automatic pipeline converts over 50,000 hours of embodied video into 21.51M trainable segments, aligning textual steps with visual state transitions. Experiments demonstrate competitive performance on embodied understanding and reasoning benchmarks, alongside general image generation and joint language-visual subgoal planning capabilities. Its joint subgoals substantially improve downstream policy performance on long-horizon manipulation tasks, while an action-generation model built on RxBrain enables effective continuous action generation on real robots without additional action pretraining.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.