Intermediate-Annotation-Free Latent Visual Reasoning via Multi-level Pseudo-Labeling
Abstract
Latent Visual Reasoning (LVR) enables Large Vision-Language Models (LVLMs) to perform intermediate reasoning in continuous representations, avoiding the need to verbalize visual information through discrete text. However, existing LVR methods either rely on costly intermediate annotations or learn latent tokens indirectly from final answer supervision, providing limited guidance on what the latent tokens should encode. We introduce Multi-level Pseudo-labeled Latent Visual Reasoning (MP-LVR), which derives explicit latent supervision directly from the pretrained LVLM's own internal signals without requiring intermediate human annotations or external teacher models. MP-LVR exploits layer-wise query-relevance attention patterns to adaptively construct multiple pseudo-labels, providing distinct supervision to individual latent tokens. In addition, we introduce Stochastic Information-Path Suppression, which reduces direct answer access to image and query tokens during training and encourages answer generation to utilize the latent tokens. Experiments demonstrate consistent improvements across diverse visual reasoning benchmarks and model configurations, while further analyses confirm that the latent tokens are actively utilized during answer generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.