Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Inference
Abstract
Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control from raw pixels. Published JEPA world models for pixel-based control (PLDM, DINO-WM, LeWM, V-JEPA-2-AC) each commit at training time to a single inference paradigm, trajectory optimisation in a learned dynamics model, so a checkpoint trained to plan cannot be read off as a policy without retraining. Deferring that choice to inference would let the constraints that bind at deployment, per-step latency and whether a goal observation can be supplied, pick the path. We present Qantara, an end-to-end JEPA of about 21M parameters whose joint training objective pairs a Brownian-bridge interpolant between consecutive clean latents on the state axis with noise-to-data flow matching on the action axis. One checkpoint serves three inference paradigms without retraining: latent planning, behaviour-cloning action sampling, and a video-inverse composition that predicts the next latent without action conditioning, then recovers the action bridging the two. Training samples the two noise levels along the edges of the (action-noise, state-noise) square, one edge per inference dispatch, so each path is trained where it is queried. Replacing that sampler with uniform sampling over the square at matched compute drops Push-T planning from 90.1 to 53.3 success rate (SR). On the LeWM control suite at a matched 10-epoch budget, Qantara averages 91.2 SR over the four environments and reaches 93.7 on OGBench-Cube, +7.7 over DINO-WM and +17.0 over LeWM, each measured against the stronger of that baseline's published score and our retrain of it. From the same weights the two goal-blind paths, which receive no goal image, reach 82.1-83.2 SR on Push-T and 70.8-72.8 on Cube at roughly 15-65x lower inference cost, and fall to near-noise on Reacher, whose released training corpus comes from a random policy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.