Heterogeneity-Aware Value Learning for Long-Term Recommendation
Abstract
Reinforcement learning (RL) provides a principled framework for optimizing long-term user engagement in recommender systems. Temporal-difference (TD) learning supports this objective by reusing trajectory segments through bootstrapped value estimation. In recommendation, sparse individual histories motivate knowledge transfer across users through a shared value function, yet heterogeneous interaction dynamics complicate this generalization. Updates induced by one user's experience may distort value estimates for others, with the resulting errors propagating through subsequent TD targets. To address this challenge, we propose Heterogeneity-Aware Value Learning (HAVL), which connects experience grouping with explicit control of shared updates. First, Heterogeneity-Aware Experience Grouping estimates expected reward, continuation probability, and successor-state features to organize experiences by their underlying interaction structure. Transfer-Constrained Value Learning then minimally adjusts proposed critic updates subject to reference TD-loss constraints defined by these groups. We characterize the grouping criterion through Bellman-backup discrepancies over a bounded affine continuation class and establish a local bound on group-level reference-loss increases. Experiments on real-world recommendation datasets demonstrate consistent cumulative-reward improvements across four RL backbones, complemented by an online evaluation in a production recommendation system.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.