acceptodds
Under review as a conference paper at ICLR 2027

Q-JEPA: Predicting Question-Relevant Visual Representations

Abstract

Acquiring visual representations that capture question-relevant regions is central to visual understanding. Existing approaches obtain such representations through explicit image manipulation, attention-guided visual selection, or autoregressive latent generation, requiring additional visual processing or intermediate generation within the language model. As the transition from a full image to a question-relevant local view can be considered as a zoom-in prediction in video, we propose Q-JEPA, which draws on V-JEPA’s video-pretrained spatiotemporal representations and predictive priors to infer question-relevant target-view representations directly from the full-image context. These representations are predicted independently of the LLM’s generation process, providing an effective and flexible visual interface with the potential to reduce inference overhead. Specifically, Q-JEPA introduces a Question-Guided Predictor that builds on the pretrained predictor with global cross-attention module to incorporate the full question context and focus cross-attention module to further associate target-view features with object- and relation-bearing contents in the question, thereby guiding prediction toward the relevant visual evidence. In addition, a three-stage training strategy progressively enables the predictor with the ability to produce question-relevant visual representations and subsequently aligns these representations with a downstream LLM for visual understanding. Evaluations show that Q-JEPA achieves 64.10% on GQA and 79.03% on PhysObjects, outperforming Qwen2.5-VL-7B by 3.72 and 19.58 percentage points, respectively, while reducing image inference latency by 22.2%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.