FlexRec: Adapting LLM-based Recommenders for Flexible Needs via Reinforcement Learning
Abstract
Modern recommender systems must adapt to dynamic, need-specific objectives, yet most are optimized for one static target. Reinforcement-learning post-training gives LLMs strong instruction-following abilities, suggesting a route to aligning them with complex recommendation goals. We study closed-set autoregressive ranking, where an LLM permutes a fixed candidate set given user context and a need instruction. Applying RL here raises two challenges: (i) an item's value changes across objectives, but sequence-level rewards collapse a ranking to one scalar and cannot identify which placements satisfy the current need; and (ii) interaction feedback is sparse and noisy, with reward reliability varying across users, items, and needs, causing inefficient and unstable updates. We propose FlexRec, a post-training RL framework combining (1) a sequential-dependence-aware item-level reward based on swaps within the remaining candidate pool and (2) critic-guided, uncertainty-aware scaling that models reward uncertainty and down-weights low-confidence signals. Across recommendation scenarios and objectives, FlexRec outperms strong baselines and improves NDCG@5 by up to 59%, Recall@5 by up to 109.4% in need-specific ranking, and Recall@5 by up to 24.1% in generalization over the base model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.