Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
Abstract
Multimodal Large Language Models (MLLMs) excel at visual understanding via text reasoning but fall short in scenarios requiring visual imagination. Existing approaches rely on predefined external toolkits or in-line image generation, whereas humans flexibly interleave visual and textual imagination within a unified mental space. Motivated by this, and by the fact that MLLMs already encode visual and textual information in the same feature space, we argue that visual tokens can be seamlessly inserted into text-token reasoning, so that visual imagination can be carried entirely by latent features. To this end, we propose Sketch-in-Latents (SkiLa), a unified multimodal reasoning paradigm that extends the auto-regressive capability of MLLMs to natively generate continuous visual embeddings, termed latent sketch tokens, as visual thoughts. SkiLa is trained in two stages: SFT explicitly grounds latent sketch tokens via a latent visual semantics reconstruction mechanism, while RL with masked GRPO on our curated SkiLa-RL-4K further implicitly refine them. SkiLa can dynamically alternates between a textual thinking mode that produces text tokens and a visual sketching mode that produces latent sketch tokens. Extensive experiments show that SkiLa achieves superior performance on vision-centric tasks and generalizes well to diverse multimodal benchmarks. Code, model and datasets will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.