HoloBrain-Discrete: Shared Action Tokenization for Cross-Embodiment Robot Learning
Abstract
Learning a single policy across robot embodiments requires an action representation that retains joint-level motion while accommodating different robot structures. Joint-graph trajectories, which record each joint's angle and link pose, expose this structure but create many correlated prediction targets. We present HoloBrain-Discrete, which couples a virtual-joint action tokenizer with a code/cue delayed autoregressive policy. Virtual joints serve as latent bottleneck nodes positioned on the robot graph, compressing trajectories into discrete codes with one encoder and decoder shared across embodiments. The policy generates these codes in groups along temporal, virtual-joint, and coarse-to-fine refinement axes, advancing one group per step rather than one code at a time. A separate prediction cue for each code gathers its context from earlier groups. Compact, structured action targets support efficient cross-embodiment pretraining and few-shot transfer. Pretraining on approximately 10,000 hours spanning 18 embodiments yields higher success than diffusion-based HoloBrain at every measured checkpoint across five evaluation settings. Cross-embodiment demonstrations also boost few-shot learning: co-training two UR5 demonstrations per task with Aloha data raises success from 6% to 39%, versus 7% to 17% for HoloBrain. Model inference is 1.9–2.1 faster than HoloBrain on LIBERO and RoboTwin.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.