Token-Depth Aggregation for Improved Text Representations
Abstract
Frozen language models are widely used as text encoders, with a trainable aggregation module that combines their token representations into one vector. Existing aggregation methods, from mean pooling to learned attention and token graphs, read the final layer and are mostly evaluated on benchmarks whose labels reflect the overall content of a sentence. Many uses instead require composing local meaning, such as a negated opinion or a chain of family relations, and aggregating it across a text, such as the majority opinion about one business among reviews of several. In this paper, we ask how the representations of a frozen language model should be combined across depth and across tokens when the label is built from such local information. We introduce a suite of composition and aggregation tasks whose labels are built from local judgments while global cues are removed, and propose Token-Depth Aggregation (), in which each token attends over its own representations at all layers before a token graph relates the fused tokens. Our experiments show that existing aggregation methods can remain close to chance level when the answer depends on tying opinions to a specific target, whereas ToDA reaches high accuracy. Across frozen encoder and decoder backbones, ToDA achieves the highest accuracy in nearly every setting of the proposed tasks. On the standard GLUE, IMDB and MTEB benchmarks, ToDA also achieves competitive or improved performance while keeping the backbone frozen and training a small number of parameters.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.