acceptodds
Under review as a conference paper at ICLR 2027

Pixel as the Common Language: Unified Multimodal Embedding via Visual Rendering

Abstract

Multimodal embedding models map heterogeneous inputs into a shared representation space, yet text-only, image-only, and composed inputs are still presented in distinct formats and follow different routes through the model. This fragmentation prevents heterogeneous tasks from jointly optimizing the same end-to-end encoding process. To address this issue, we propose unification before encoding and instantiate this paradigm in RenderEmbed. The framework uses Visual Substrate Rendering to convert heterogeneous inputs into a common visual substrate and then encodes the resulting visual-token sequence with a VLM. To retain both broader context and fine-grained cues in compact retrieval representations, it applies Progressive Hierarchical Anchor Encoding by interleaving anchors throughout the visual-token sequence and fusing their representations across layers. It further employs Interventional Rendering Learning to distinguish retrieval semantics from presentation variations through complementary content and rendering interventions. Extensive experiments on MMEB and ViDoRe v2 show that RenderEmbed outperforms prior state-of-the-art models at the same scale and exhibits strong out-of-domain generalization. Compared with MetaEmbed, RenderEmbed improves MMEB Overall Precision@1 by 1.7 points, with a larger 2.1-point gain on out-of-domain tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.