acceptodds
Under review as a conference paper at ICLR 2027

ALICE: Aligning Latent Reasoning with Chain-of-Thought for Multimodal-LLM Embeddings

Abstract

Recent multimodal large language model (MLLM) embedders show potential to improve retrieval by generating explicit chain-of-thought (CoT) rationales before the final representation. However, this introduces substantial inference latency and may incur information loss due to discrete autoregressive generation. In this paper, we investigate bringing latent reasoning into multimodal embedding. We first show that latent tokens trained with curriculum-based next-token prediction tend to collapse to predicting a narrow set of CoT-initial tokens, suggesting a bias toward local token prediction over holistic reasoning semantics. Based on this observation, we propose **ALICE** (**A**ligning **L**atent Reason**i**ng with **C**hain-of-Thought **E**mbeddings), a direct *hidden-state to embedding-space* alignment approach: we insert a short block of `<latent>` tokens into the query and train their hidden states to match per-step CoT anchors from a separately trained multi-granularity reference embedder, followed by a free-form contrastive stage that adapts the latents for retrieval. ALICE thus requires a single forward pass at inference, without autoregressive generation. On MMEB, ALICE achieves strong results across the 2B–8B scale, reaching 74.8 with Qwen2-VL-7B, with minimal additional latency. Analysis further reveals that the learned latents exhibit sequential structure, encode instance-specific information, and become increasingly aligned with the final embedding, suggesting reasoning-like behavior in latent space.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.