On the Vulnerability of Semantic IDs in Generative Recommendation
Abstract
Generative recommenders (GenRec) are increasingly adopted on real-world platforms. Semantic IDs (SIDs) are widely used in GenRec to map item descriptions to discrete tokens for autoregressive recommendation. Editing a listing can change its SID while the recommender remains fixed, creating a new attack surface for targeted item promotion. We introduce SIDAttack, which appends a short suffix to a target item's description while treating the victim SID tokenizer and recommender as black boxes. It searches using gradients from available text encoders and public-item anchors selected from publicly visible popular items, without victim recommendation feedback. Across product, news, and short-video catalogs and two SID tokenizers, SIDAttack produces the largest mean exposure gains among the evaluated baselines when the search and victim encoders match. Although the gain is smaller when the victim encoder is held out, SIDAttack still increases exposure relative to the unattacked baseline. We also find that residual-quantization margins strongly predict which items undergo SID changes. Guided by this finding, we propose margin-regularized codebook construction, which reduces attack-induced exposure while largely preserving clean recommendation utility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.