CrystaL: Seeing Less to Learn Better Visual Latents
Abstract
Learning visual latent states that effectively support answer generation remains challenging for multimodal large language models (MLLMs). Existing approaches often rely on explicit visual supervision or staged restrictions on image access, introducing dependencies on predefined visual targets or multi-stage training. We introduce CrystaL, a single-stage dual-path framework that learns task-relevant visual latent states through their utility for answer generation. An intact path derives latent states from the original image, while a parameter-shared corrupted path leverages the transferred states to predict the same answer under degraded visual input. Joint answer supervision and cross-path consistency encourage the latent states to preserve information useful for prediction when direct visual evidence is weakened. CrystaL requires neither external visual feature targets nor intermediate visual annotations, and retains only the intact path at inference. Across nine benchmarks, CrystaL achieves relative gains of 6.0% and 4.6% over Qwen2.5-VL-7B and Qwen3-VL-4B,respectively, supporting the effectiveness of the proposed dual-path training framework in learning visual latents that support answer generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.