acceptodds
Under review as a conference paper at ICLR 2027

DG-HCA: Dual-Granularity Hierarchical Credit Assignment for Generative Recommendation

Abstract

Generative recommendation with explicit reasoning typically relies on sequence-level reinforcement learning, where a single outcome-level advantage is broadcast to all tokens in a rollout. However, for hierarchical Semantic IDs (SIDs), such coarse-grained supervision can induce conflicting credit signals: incorrect suffix decisions may be reinforced by positive advantages, while correct prefix decisions may be suppressed by negative ones. To address this issue, we propose DG-HCA, a Dual-Granularity Hierarchical Credit Assignment framework for reasoning-based generative recommendation. DG-HCA grounds policy optimization in retrieval feedback from Constrained Beam Search and decomposes credit at two complementary granularities. At the reasoning-trajectory level, Ranking-Driven Reasoning Credit evaluates trajectories according to their retrieval quality. At the SID-action level, Target-Path SID Credit performs relative credit assignment over hierarchical decisions along the ground-truth path, while Target-Path Recovery provides additional supervision when Beam Search deviates from the target SID. Extensive experiments on three public Amazon datasets demonstrate that DG-HCA achieves competitive recommendation performance and obtains the best NDCG@10 on all three datasets. Ablation studies further verify the effectiveness of SID-level credit assignment and Target-Path Recovery. These results highlight the importance of aligning reinforcement-learning credit with the structured generation process in reasoning-based generative recommendation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.