acceptodds
Under review as a conference paper at ICLR 2027

CrystaL: Seeing Less to Learn Better Visual Latents

Abstract

Learning visual latent states that effectively support answer generation remains challenging for multimodal large language models (MLLMs). Existing approaches often rely on explicit visual supervision or staged restrictions on image access, introducing dependencies on predefined visual targets or multi-stage training. We introduce CrystaL, a single-stage dual-path framework that learns task-relevant visual latent states through their utility for answer generation. An intact path derives latent states from the original image, while a parameter-shared corrupted path leverages the transferred states to predict the same answer under degraded visual input. Joint answer supervision and cross-path consistency encourage the latent states to preserve information useful for prediction when direct visual evidence is weakened. CrystaL requires neither external visual feature targets nor intermediate visual annotations, and retains only the intact path at inference. Across nine benchmarks, CrystaL achieves relative gains of 6.0% and 4.6% over Qwen2.5-VL-7B and Qwen3-VL-4B,respectively, supporting the effectiveness of the proposed dual-path training framework in learning visual latents that support answer generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.