acceptodds
Under review as a conference paper at ICLR 2027

Do Reasoning Models Enhance Embedding Models?

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) makes large language models better problem solvers, but does it also make their representations better foundations for text embeddings? We compare embedding models initialized from matched base backbones and backbones tuned with RLVR, using identical recipes for contrastive learning. Across MTEB Multilingual, MTEB Code, and BRIGHT, RLVR initialization provides no consistent advantage: the absolute paired mean differences reach at most 0.67 points. To explain this result, we introduce Hierarchical Representation Similarity Analysis (HRSA), which separates three properties often conflated by a single similarity measure: feature correspondence at the coordinate level, global and local manifold geometry, and functional readouts relevant to downstream tasks. HRSA shows that RLVR can reorganize local manifold geometry and, under prolonged training, shift the coordinate basis, while largely preserving global manifold geometry and linear readout structure. When the base and reasoning backbones are subsequently trained for embeddings, their shared contrastive objective rapidly pulls them toward the same global and functional organization, although differences in local manifold geometry persist. We call this phenomenon Manifold Realignment. In contrast, supervised fine tuning more strongly disrupts coordinate and geometric alignment. These findings suggest that RLVR primarily changes how a model traverses and locally organizes an existing semantic landscape rather than constructing a globally superior one. Consequently, stronger reasoning behavior does not necessarily produce better embedding models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.