Codes for Representation, Items for Prediction: Rethinking the Learning Interface of Semantic-ID Recommendation
Abstract
Semantic IDs (SIDs) represent items through compositions of shared discrete codes, yet generative recommenders typically also use these codes as autoregressive prediction targets. We observe that, under the same backbone configuration, direct item prediction can outperform autoregressive SID prediction, motivating the hypothesis that representation sharing and prediction factorization need not be coupled. To examine this hypothesis, we analyze what changes when a complete-item decision is factorized into an autoregressive SID path. At the distribution level, a locally normalized legal SID tree can represent any strictly positive catalog distribution, while controls over invalid-continuation competition do not explain the observed ranking gap. This directs attention to the supervision structure: under teacher forcing, SID prediction directly supervises branch logits along the target path, progressively refining nested target subtrees, whereas non-target subtrees compete through shared entry branches without direct supervision of their internal conditionals from the current example. Progressively controlled comparisons show that direct item scoring remains stronger under increasingly matched designs and exact catalog ranking. Based on this decoupling principle, we propose Next-item Representation Prediction (NRP), which composes shared SID embeddings into complete-item representations and directly optimizes item-level scores without an item-specific embedding table. Across three public datasets, NRP achieves competitive performance against strong sequential and generative baselines and improves transductive cold-start recommendation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.