acceptodds
Under review as a conference paper at ICLR 2027

Contrastive Label-Embedding Alignment: Adapting Dual Encoders for Text Classification from Label Descriptions

Abstract

Dual-encoder text embedding models provide efficient, reusable representations, but adapting them to a target classification task typically requires labeled documents to fit a classifier head. Alternatives that avoid labeled documents, such as NLI cross-encoders and prompted LLMs, process each document jointly with its candidate labels and therefore cannot cache document representations. We introduce __contrastive label-embedding alignment__ (CLEA), which reaches a point on this supervision-cost / inference-cost frontier that, to our knowledge, no prior method occupies: no labeled documents, no task-specific head, and fully pre-encodable dual-encoder inference. CLEA lightly fine-tunes a shared encoder using only label verbalizers and a small set of natural-language descriptions per label, aligning each verbalizer with its description set through a symmetric multi-positive contrastive objective. Across four benchmarks spanning topic, sentiment, intent, and emotion classification and ten encoders from 22M to 600M parameters, CLEA improves macro-F by __+0.09__ on average over untuned zero-shot embeddings, and on every benchmark the best CLEA-adapted encoder matches or exceeds NLI cross-encoders, a reranker, and prompted LLMs while retaining pre-encodable inference. Five descriptions per label suffice, and LLM-generated descriptions recover most of the gains of hand-written ones. In budget-matched experiments, CLEA outperforms linear probes until several dozen labeled examples per class are available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.