acceptodds
Under review as a conference paper at ICLR 2027

Aligning Where It Matters: Critical-Branch Preference Optimization for LLM-based Generative Recommendation

Abstract

Generative recommendation increasingly represents items with hierarchical Semantic IDs (SIDs) and formulates next-item prediction as autoregressive token generation. While pretrained large language models (LLMs) have shown strong potential for SID-based recommendation, how they exploit the internal structure of SIDs remains insufficiently understood. In this work, we conduct a fine-grained analysis of hierarchical SIDs in LLM-based generative recommendation. We find that the advantage of pretrained LLMs is highly concentrated at early SID levels, which are both sufficient and necessary for recovering item semantics and are more critical for preserving behavior-relevant information. This observation motivates us to explicitly optimize early SID decisions through preference learning. However, we show that standard Direct Preference Optimization (DPO) is mismatched with autoregressive SID generation: suffix tokens introduce both margin-scale and gradient-direction interference, such that improvement of the critical decision cannot be guaranteed. To address this issue, we propose Critical-Branch Direct Preference Optimization (CB-DPO), which applies preference supervision directly at the first-divergence position of each preference pair, providing a cleaner optimization signal and guaranteeing an increase in the critical margin. Experiments on KuaiRec and KuaiRand-1K demonstrate the effectiveness of CB-DPO across both LLM-based and LLM-free generative recommendation models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.