Same Rankings, Different Semantic IDs: Optimization Non-Invariance in Tokenization
Abstract
Semantic-ID-based generative recommendation relies on tokenizers such as RQ-VAE to convert continuous item representations into discrete codes. Although scale and distribution are known to affect codebook collapse, their effect on the full recommendation pipeline is harder to isolate when the upstream representation also changes. We study input rescaling that preserves every pairwise recommendation-score difference. In our main study with 512-dimensional SASRec embeddings, the same training procedure collapses to a single base SID on native inputs, whereas rescaling produces broad codebook use and substantially better downstream recommendation. An exact tokenizer reparameterization shows that the same non-collapsed code assignments remain representable at native scale, separating representational capacity from training outcome. At these mapped states, reconstruction loss and its encoder-output gradient scale quadratically, whereas latent quantization terms remain unchanged. This analysis yields a simple reconstruction-loss calibration that keeps the native embeddings fixed. Controlled experiments from matched starting points show that the calibration changes encoder optimization and recovers broad code use. Across three Amazon datasets, calibration improves Hit@10 and NDCG@10 in all nine paired runs, reaching accuracy close to direct input rescaling. These findings show why representation comparisons should account for numerical scale and tokenizer optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.