Eyes in the Update? Learning Where Rewards Land in Multimodal Reasoning
Abstract
Multimodal post-training has moved from broad instruction tuning toward reinforcement learning for image and video reasoning, where outcome rewards can improve final-answer accuracy. Yet these rewards are assigned to complete generated trajectories, while visual evidence is used only at sparse and uneven positions inside the response. A growing line of work has therefore begun to make multimodal RL more token-sensitive through grounding signals or external token-level weighting. We ask a more specific question: can the allocation of the same trajectory-level reward be inferred from the model's own multimodal computation, rather than from an auxiliary annotator, perception module, or process label? We propose Gaze-RL, a representation-driven policy optimization method that redistributes completion-level advantage through a latent reward landing field. The field is computed from internal cross-modal geometry and local predictive uncertainty, assigning larger update weight to positions that are both visually anchored and decision-bearing while remaining mean-one, so the total credit of each rollout is preserved and only its positionwise allocation changes. Gaze-RL requires no process supervision, auxiliary perception module, or architectural modification. Trained on a mixed image-video RL dataset, Gaze-RL improves Qwen2.5-VL and InternVL3.5 across broad image and video benchmarks. It also yields stronger evidence sensitivity and shorter, more concentrated reasoning dynamics. These results suggest that multimodal RL benefits from learning how existing rewards should land inside the reasoning trajectory.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.