acceptodds
Under review as a conference paper at ICLR 2027

JEVA: Joint Encoders for Variational Action World Models

Abstract

Learning world models for robotic control typically requires large amounts of action-labeled data, which is expensive to collect, whereas action-free demonstrations such as video are abundant. Latent action models address this imbalance by inferring the actions that explain observed transitions, but training the inverse dynamics model (IDM) jointly with the forward model is unstable. Existing methods therefore fix the latent action space before training the world model, or add a separate inverse-dynamics warm-up stage. We introduce JEVA, a latent-action world model that co-trains an IDM and a forward world model end-to-end from scratch in a single stage, on a pool of trajectories in which only a small fraction of transitions carry action labels. JEVA treats the action between two visual states as a continuous Gaussian latent that can be planned over and decoded into continuous controls. Variance-reduced routing (VRR) is what makes single-stage training work. On labeled transitions, where the action decoder keeps posterior forward predictor on the posterior mean instead. This removes reparameterization noise from the predictor gradient at the cost of a small bias, and it also removes a penalty that otherwise drives latent actions toward collapse: with a warm-up stage and no VRR every latent-action dimension collapses. With 5% of episodes labeled, a policy trained in JEVA's latent action space matches or exceeds every baseline trained with 20% of labels on OGBench Cube-Double and Cube-Triple. Under CEM planning, JEVA leads the strongest world model baseline by 11-22 points on all three Cube tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.