Train Fine, Serve Coarse: Tokenizer-Agnostic Item-Level Alignment for Generative Recommendation
Abstract
Generative recommendation models commonly undergo post-training to align their outputs with application-specific objectives such as relevance or user preference. The rewards used for alignment are typically defined on individual items, whereas policy optimization operates on the semantic IDs that encode these items. This creates a mismatch between the granularity of reward evaluation and policy optimization when multiple items share a semantic ID, or node. Items within the same node can receive different rewards for the same input, yet their rewards affect a single shared node likelihood, limiting the policy's ability to learn from these distinctions. We propose Depth-Extended Alignment (DEA), which introduces a training-time refinement of the semantic-ID hierarchy to enable policy optimization at the granularity of individual items. We derive a hierarchical alignment objective that accounts for reward differences across nodes and among items within each node, while shared representations enable item-specific feedback to update the generative model. At inference, DEA preserves the original node-level decoding process without additional item-level computation. Offline experiments on public and real-world industrial datasets demonstrate substantial performance improvements with both RQ-Kmeans and RQ-VAE, two widely used tokenization methods. Validated through rigorous online A/B testing, DEA now operates in full production on a large-scale e-commerce platform serving hundreds of millions of active users.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.