An LLM Hiddens Several Semantic Experts: Probing Task-Specific Attention Values for Text Embedding
Abstract
Different text embedding tasks require different notions of semantic similarity, motivating representations that adapt to the task. In this paper, we investigate whether its attention head modules already contain task-specific semantic signals. Specifically, we introduce a training-free framework that probes and ranks individual modules using augmented text pairs or downstream reference data, then constructs compact task-specific embeddings from their attention values. Two complementary augmentations enable selection using unlabeled data; task instructions can also select and combine dataset-specific experts. Probing is performed offline and routing is decided once per task, while each text requires only one frozen-model forward pass. Across 14 MTEB tasks, eight selected modules yield 1024-dimensional embeddings that exceed the strongest prompt-based ensemble by 4.66 and 5.08 average points on Llama3-8B and Qwen3-8B, respectively. Each text requires one forward pass; the ensemble uses eight. Dynamic routing achieves the highest mean at this dimension on both models. On a trained LLM2Vec embedder, selected modules surpass its native 4096-dimensional representation with either 2048 or 4096 dimensions. These results show that offline task-specific probing can turn existing dense models into flexible, compact embedders.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.