What Makes Semantic IDs Learnable? A Spectral View of Output Factorization.
Abstract
Generative recommender systems commonly use semantic IDs (SIDs) to represent items as sequences of tokens that are in turn predicted autoregressively. When each item has a unique code, different SID mappings preserve the same item information but decompose item prediction into different sequences of token predictions. We study this output factorization: what properties of a code make its successive predictions easier for a model with restricted output rank? We describe atomic and SID mappings as filtrations and derive spectral approximation bounds for the best predictor in a specified class. These bounds motivate the gain-weighted spectral tail (GWT), a mapping-level proxy for learnability at a given head rank. Controlled synthetic experiments isolate spectral concentration at matched predictive gain and numerical rank, showing its consequences for learning. We then assess whether GWT remains useful after training rank-restricted SASRec on three Amazon datasets. With atomic histories fixed, GWT exhibits strong correlation with test negative log-likelihood (NLL) across seven SID mappings, including RQ-VAE and RQ-Kmeans. We also compare GWT with common SID diagnostics, including codebook utilization and the per-level token count Gini coefficient. These initial results support GWT as a principled proxy metric for SID learnability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.