Does Positional Encoding Matter for Semantic ID-based Generative Recommendation?
Abstract
Semantic IDs have emerged as an effective representation for generative recommendation, where each item is represented by multiple discrete codes. These codes are typically serialized along the user interaction sequence into a flattened token sequence for generative modeling. However, such a formulation couples two structurally different sources of positional variation: the sequential order of user interactions and the hierarchical order induced by residual quantization. We argue that these two structures should be modeled on their respective positional axes. Based on this insight, we propose SimPos, a simple structure-aware positional modeling method that replaces the Transformer backbone's positional encoding with separate interaction-level and residual-level attention biases, requiring no additional learnable parameters. Experiments on four real-world recommendation datasets with different backbones and residual tokenizers show that SimPos consistently outperforms existing positional encoding methods, with relative gains of up to +25.25% in Recall@5. These results indicate that positional encoding matters in semantic ID-based generative recommendation, and that directly inheriting it from the Transformer backbone is suboptimal. Our code and data are available at https://anonymous.4open.science/r/SimPos-8CFC/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.