acceptodds
Under review as a conference paper at ICLR 2027

TraViS: Learning Latent Visual Trajectories with Context-Adaptive Visual Targeting

Abstract

Multimodal large language models (MLLMs) often rely on localized visual evidence, yet existing supervision provides limited guidance on what intermediate latent states should visually predict. We introduce TraViS, a framework for learning latent visual trajectories with context-adaptive visual targets. Instead of using predefined intermediate supervision, TraViS enables each latent state to identify task-relevant visual evidence from the evolving multimodal context. Specifically, we propose a state-conditioned target construction mechanism, where a frozen visual teacher encodes the selected region into continuous visual targets, and a target feedback mechanism that injects these targets into subsequent latent computation during training. At inference, TraViS removes the teacher and performs self-feedback latent rollout with only a single global image encoding, avoiding additional visual processing. Extensive experiments on multiple backbones and eight multimodal reasoning benchmarks demonstrate that TraViS consistently improves performance, achieving average gains of over 2.5% while maintaining a favorable accuracy-efficiency trade-off.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.