InnerJEPA: Masked Visual Latent Prediction for Video Understanding and Generative Transfer
Abstract
Video multimodal large language models (MLLMs) are typically instruction-tuned with answer-token supervision, leaving visual prediction implicit. We introduce InnerJEPA, a lightweight auxiliary objective that turns an existing video MLLM into a masked visual latent predictor. Its frozen vision encoder supplies clean target features, while its causal language model predicts masked visual representations through a small projection head. Masking is applied after vision encoding, allowing inputs and targets to share a single encoder pass. The mask embedding and prediction head are discarded at inference. However, naive cosine regression admits a shortcut in the anisotropic target space: predictions become nearly parallel despite decreasing reconstruction loss. InnerJEPA combines target centring, leading-direction removal, and variance–covariance regularisation to counter this degeneracy. In our representation diagnostic, these changes increase prediction effective rank from 14 to 332 and reduce mean pairwise cosine similarity from 0.908 to 0.160. In a controlled Qwen3.5-0.8B comparison, InnerJEPA improves schedule-matched supervised fine-tuning by 1.07, 1.72, 2.45, and 3.09 percentage points on MVBench, TempCompass, NExT-QA, and STAR, respectively, averaging a 2.09-point gain. It also improves four of five temporal-grounding benchmarks, with a 2.34-point gain in mean mIoU. Beyond video understanding, a matched text-to-image transfer experiment using identical alignment-MLP and DiT-LoRA adaptation settings raises CLIP score from 0.2329 to 0.2449 on 553 images generated from GenEval prompts. A blinded pilot on a fixed 20-prompt subset favours InnerJEPA over SFT on 14 prompts, with five ties. These results support collapse-resistant masked latent prediction as useful auxiliary supervision for video understanding and provide evidence of transfer to image generation without adding inference-time components to the video MLLM.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.