InnerEye: Giving Multimodal Agents a Pre-Action Visual State
Abstract
Multimodal agents increasingly operate in interactive environments by interpreting visual inputs, maintaining interaction context, and generating executable actions. In many such agents, the explicit decision trajectory is progressively represented through autoregressive token sequences, while visual features remain available within the underlying multimodal model. However, agent decisions are inherently step-dependent: as interaction history evolves, the visual evidence relevant to the next action may change even when the visual input itself remains unchanged. This raises an agent-level question beyond visual perception itself: how should visual information be represented for the current decision before it is expressed as an action? We introduce InnerEye, a multimodal agent formulation that explicitly constructs a step-specific pre-action visual state before action generation. At each step, this transient state is freshly formed from the task instruction, visual input, and current interaction history, providing a latent representation that supports the upcoming action. InnerEye progressively trains this state from visual grounding toward supporting the upcoming action. Stage 1 performs visually grounded supervised fine-tuning by aligning latent hidden states with features from a frozen vision encoder while learning expert actions. Stage 2 further optimizes the grounded state through reinforcement learning with Online Latent Action Distillation, which transfers the pre-action visual states from a full-vision branch to an auxiliary branch without direct access to image tokens on successful rollouts, providing an action-oriented training signal without altering the standard full-vision inference process. Across two Qwen3-VL scales and three benchmarks spanning multi-step visual tool use and single-step visual action prediction, InnerEye consistently improves over both standard post-training baselines and a latent visual reasoning baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.