VISER: Learning to See What Matters for Efficient Visual Reasoning
Abstract
Recent “Thinking with Images” methods have improved visual reasoning by explicitly inspecting image regions or reasoning with latent visual representations. However, many of these methods introduce additional visual interactions or modify the base model's inference procedure. How can models learn to better attend to and use visual evidence while maintaining efficient inference? We propose VISER, a two-stage post-training framework that combines visual-token evidence supervision with reinforcement learning to guide evidence use in answer generation. The first stage jointly optimizes evidence prediction at visual-token positions and answer generation. The second stage derives contrastive rewards from the model's responses to original, evidence-corrupted, and background-perturbed images, combining them with task-correctness rewards for reinforcement learning. These rewards encourage reliance on visual evidence and robustness to irrelevant background perturbations. All additional supervision is confined to training; VISER preserves the base model's inference procedure without additional region cropping, visual tool calls, latent visual reasoning tokens, or multi-turn visual interactions. Experiments demonstrate that VISER achieves the best overall performance among the compared methods across five image VQA benchmarks while maintaining inference latency close to that of the base model. Under the same evaluation setup on the full VStar benchmark, VISER incurs approximately 4% latency overhead over the base model and achieves approximately 2.4×, 58×, and 9.3× speedups over TreeVGR, DeepEyes, and DeepEyesV2, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.