RADAR: Radial Attribute Distillation for Augmented Representations
Abstract
Dense embeddings primarily encode similarity through the vector direction, while vector magnitude is often normalized away or lacks an explicit semantic role. We introduce RADAR (Radial Attribute Distillation for Augmented Representations), a general framework for augmenting pretrained embeddings with latent attributes without disrupting their underlying semantic geometry. RADAR learns from multiple heterogeneous and noisy proxy signals to identify a latent attribute that is both supported by the proxies and predictable from the original representation. It then encodes the predicted attribute through bounded radial scaling: embedding direction—and hence the original semantic representation—is preserved, while vector norm carries complementary information that can be exploited directly by inner-product retrieval. The framework is domain-agnostic and does not require retraining the underlying embedding model or observing proxy signals at inference time. We evaluate RADAR using scientific literature retrieval as a challenging testbed, where citations and venue-related signals provide imperfect proxies for scholarly impact. Across four retrieval benchmarks, augmenting Qwen3-Embedding-8B with RADAR improves average nDCG@10 from 45.16 to 46.02 and Recall@50 from 50.01 to 51.31, while retaining single-stage vector retrieval. The distilled attribute also exhibits stronger associations with held-out peer-review judgments and prospective citation outcomes than direct proxy prediction and conventional factor-based alternatives on most measures. Our results suggest that radial degrees of freedom provide a simple mechanism for incorporating complementary latent attributes into existing embedding spaces while preserving their semantic structure and computational advantages.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.