acceptodds
Under review as a conference paper at ICLR 2027

The Update Bottleneck: Re-examining Joint Training for Semantic-ID-Based Generative Recommendation

Abstract

Semantic ID based Generative Retrieval models map items to discrete Semantic IDs (SIDs) and train a sequence model over the codes. Several recent methods train the tokenizer jointly with the recommender and report 9 to 14% gains over two-stage pipelines. We re-examine this comparison under matched conditions. Joint training improves Recall@10 on Beauty without a statistically significant gain against a two-stage method, while Recall@10 falls by approximately 8–16% on the other four datasets. We propose the update bottleneck hypothesis to explain these inconsistent benefits: tokenizer updates may leave discrete IDs unchanged, while frequent ID changes can disrupt the recommender's learning. Our theory formalizes the limits on how tokenizer updates affect a recommender through discrete IDs. Controlled experiments support the update bottleneck and show why gradient flow alone does not ensure useful representation learning in semantic-ID-based generative recommendation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.