acceptodds
Under review as a conference paper at ICLR 2027

Why Only the Final Thought? Hierarchical Aggregation for Multimodal Retrieval

Abstract

Multimodal large language models (MLLMs) have emerged as powerful backbones for learning embeddings for multimodal retrieval. Recent methods incorporate chain-of-thought (CoT) reasoning into embedding generation and further employ adaptive routing mechanisms to selectively invoke reasoning and control its computational cost. However, different queries may benefit from embeddings formed with different reasoning-token budgets and at different model depths, whereas existing methods typically rely on a single final embedding. To address this limitation, we introduce a dual-aggregator framework that exploits these embeddings, adaptively combining them for each query and aggregating their matching evidence to rank candidates. Experiments on MMEB-V2 validate the benefits of integrating information across reasoning-token budgets and model depths, with our framework achieving strong overall retrieval performance with less than 1% additional parameters over the backbone. Further analytical experiments demonstrate that our framework adapts its aggregation to individual queries, effectively integrating information across budgets and depths to improve retrieval performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.