acceptodds
Under review as a conference paper at ICLR 2027

Juno: Taming Predictive Latents for Vision-Language-Action Models

Abstract

Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision–language–action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: mismatch with embodiment-specific control, interference with action learning, and teacher miscalibration under distribution shifts. We introduce **_Juno_**, a unified framework built around one action-conditioned JEPA that serves as a control-aligned representation backbone, a predictive teacher, and an adaptable dynamics model. During **pretraining**, we train it on embodiment-matched trajectories and use a dynamic CLS loss to transfer motion-weighted patch dynamics to a compact global state. During **policy learning**, we fuse current-frame JEPA patches into VLA perception and use a decoupled reasoning branch with separate transformation parameters to distill future latent states for action generation. During **deployment**, we adapt the world model on all observed transitions, including failed rollouts, freeze the adapted teacher, and re-align the policy on verified executions using LoRA adapters and a trainable action head, without expert corrections or task rewards. On SimplerEnv, **_Juno_** raises average success from to over Qwen3GR00T, the strongest baseline, and test-time adaptation further reaches ; on a real robot, it retains – success under background, height, and object shifts where the base policy collapses to .

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.