acceptodds
Under review as a conference paper at ICLR 2027

IDShare: Learned Sharing as Regularization for Long-Tailed ID Embeddings

Abstract

Recommendation models often use ID features with many distinct values. Features such as user IDs and ad IDs can have long-tailed frequency distributions. Long-tailed IDs have few training examples, making their embeddings difficult to learn reliably. We introduce IDShare, a simple end-to-end method that learns which IDs share an embedding. Each shared embedding is updated using training examples from all IDs assigned to it. We find that this learned sharing provides effective regularization for long-tailed IDs. On two recommendation benchmarks, IDShareachieves higher mean AUC than independent embeddings with tuned global L2, without using that penalty. Further analysis shows that IDShare improves predictions for long-tailed IDs. Its performance remains stable across different numbers of shared embeddings. It also retains its advantage in deeper models. These findings support learned sharing as a simple and stable default for training long-tailed ID embeddings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.