One State, One Roll: Unified Physical Latent Flow for Robot Manipulation
Abstract
Embodied foundation models are increasingly moving toward multimodal prediction, jointly covering actions, proprioceptive states, and visual observations. Yet these modalities are usually generated through separate modality-specific pathways. Such separation makes cross-modal physical consistency an additional coordination problem: different pathways can predict incompatible futures, while each modality is evolved without a shared representation of the physical transition. Moreover, most models generate every future chunk from random noise, leaving the local continuity of an already realized physical trajectory largely unused. We introduce OneRoll, a unified physical latent flow for robot manipulation. OneRoll linearly fuses the three modalities into one physical token and recovers them through linear projections. The token's velocity on the physical manifold is thus the sum of their linearly mapped derivatives, directly supervising a shared vector field whose integration extends the realized trajectory into the future. Experiments on both simulated and real-robot tasks show encouraging performance, suggesting that unified physical evolution is a promising foundation architecture for embodied models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.