Branch-JEPA: Learning Dynamics-Aware State Representations through Dual-Branch Patch Routing
Abstract
Predicting future embeddings offers a compact route to learning visual world models, but a well-dispersed latent space need not expose the state changes that distinguish actions. We present Branch-JEPA, a representation-level extension of an end-to-end joint-embedding world model. Two independent spatial routers aggregate shared Patch features into State and Act residuals, which augment rather than replace a global CLS pathway. A contrastive action-alignment objective supervises changes in the fused Act coordinates, while one shared Predictor models the complete representation. The design preserves latent prediction, global anticollapse regularization and goal-conditioned planning without a pretrained visual backbone or a denoising frontend. Across PushT, TwoRoom, Reacher and Cube, task-averaged planning success increases from 83.25% to 88.25% on clean observations and from 35.50% to 59.75% under Gaussian corruption. Direct actionregression controls show that action supervision and Patch access account for much of the gain. Paired intervals resolve smaller improvements from separate routing and Act-only alignment under Gaussian noise; their clean intervals include zero. Frozen readouts, interventions and success-aligned candidate ranking connect these behavioral results to relative specialization within the shared representation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.