JACT: Learning to Act Through Frozen JEPA Dynamics
Abstract
Visual world models can guide control through online planning, but repeated pre- diction and search add deployment cost. We propose JACT (JEPA-guided Action- Consequence Training), which moves the use of predictive dynamics into offline policy training. A direct policy acts recursively inside a frozen joint-embedding predictive architecture (JEPA), conditions subsequent decisions on predicted con- sequences, and learns to align the resulting trajectory with observed future em- beddings. Gradients pass through the fixed dynamics into the actor, while local supervision anchors its actions. Deployment uses only the encoder and policy, with no predictive rollout, search, or retrieval. Across 11 visual control settings, JACT combines broad task coverage with efficient direct execution. It achieves 97.25% average success on the four-task comparison, the highest among the listed methods with complete coverage. Controlled ablations and repeated actor training support the value of recursive consequence supervision, while identifying task- dependent limits. JACT offers a simple route from pretrained JEPA dynamics to direct visual control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.