VeGo: Visual Evidence-Gated On-Policy Learning for Vision-Language Models
Abstract
Multimodal large language models still struggle with fine-grained visual understanding, particularly when answering questions that depend on small or inconspicuous image regions. Existing remedies either invoke cropping tools at inference, incurring considerable overhead from repeated tool calls and visual processing, or internalize region-focused perception through on-policy self-distillation, where a teacher observes a question-relevant crop while the student sees the full image. However, such distillation treats every generated token uniformly, although privileged visual evidence is informative for only a small subset of content-bearing tokens. Moreover, responses often continue unnecessarily after the answer has already been determined, resulting in redundant generation and inefficient inference. To address these limitations, we propose VeGo, which contrasts the teacher's next-token predictions under an evidence-rich crop and a degraded, evidence-blind view to estimate each token's reliance on fine-grained visual evidence. This enables distillation to prioritize evidence-sensitive tokens while preserving the overall loss magnitude. To reduce redundant generation, VeGo also identifies when the necessary visual evidence has been sufficiently encoded in the generated text and applies a length reward only to the subsequent region, leaving the evidence-writing prefix unaffected. Experiments across V, ZoomBench, Visual Probe, HR-Bench 4K and HR-Bench 8K show that VeGo achieves the highest average accuracy among open-source, agentic, and distillation baselines while producing the shortest responses.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.