acceptodds
Under review as a conference paper at ICLR 2027

Projector-JEPA: A Gated DeltaNet Two-Expert Residual Projector Bridging Video-LLMs to World-Dynamics Understanding

Abstract

Recent Video-LLMs have made rapid progress by combining powerful vision encoders with large language models. However, most models rely on a projector that simply maps frame-level visual features into the LLM input space, without explicitly modeling local changes between frames or long-range state transitions. This limitation can become a bottleneck in video understanding tasks that require temporal structure, such as motion reasoning, event ordering, and state change recognition. We propose Projector-JEPA, a two-expert residual projector that improves world-dynamics understanding in Video-LLMs by explicitly modeling local motion and global state transitions, while keeping the vision encoder, base projector, and LLM frozen. Projector-JEPA augments the output of the base projector with two Gated DeltaNet-based residual experts. The Local Expert processes a temporal sequence at each spatial position, while the State Expert processes spatially pooled frame summaries; both retain recurrent history. To supervise these complementary spatial granularities, we introduce JEPA-inspired auxiliary objectives that combine a local correspondence loss with a masked state reconstruction loss. We evaluate Projector-JEPA on seven Video-LLMs from the PLM, VideoLLaMA3, and InternVL3 families across five video understanding benchmarks. Projector-JEPA consistently improves performance across all models and benchmarks, with larger gains on tasks involving state changes, memory, and motion direction. These results suggest that the projector in Video-LLMs should be viewed not merely as a modality alignment module, but as a key component for conveying dynamic visual information to the LLM.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.