acceptodds
Under review as a conference paper at ICLR 2027

Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs

Abstract

Multimodal Large Language Models (MLLMs) excel at visual understanding via text reasoning but fall short in scenarios requiring visual imagination. Existing approaches rely on predefined external toolkits or in-line image generation, whereas humans flexibly interleave visual and textual imagination within a unified mental space. Motivated by this, and by the fact that MLLMs already encode visual and textual information in the same feature space, we argue that visual tokens can be seamlessly inserted into text-token reasoning, so that visual imagination can be carried entirely by latent features. To this end, we propose Sketch-in-Latents (SkiLa), a unified multimodal reasoning paradigm that extends the auto-regressive capability of MLLMs to natively generate continuous visual embeddings, termed latent sketch tokens, as visual thoughts. SkiLa is trained in two stages: SFT explicitly grounds latent sketch tokens via a latent visual semantics reconstruction mechanism, while RL with masked GRPO on our curated SkiLa-RL-4K further implicitly refine them. SkiLa can dynamically alternates between a textual thinking mode that produces text tokens and a visual sketching mode that produces latent sketch tokens. Extensive experiments show that SkiLa achieves superior performance on vision-centric tasks and generalizes well to diverse multimodal benchmarks. Code, model and datasets will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.