acceptodds
Under review as a conference paper at ICLR 2027

SimRec: Learning to Search, Ask, and Recommend from Simulated Interaction

Abstract

Recommendation agents built on large language models can search a catalog, inspect items, ask the user, and revise a recommendation after the user rejects it. In practice, however, these behaviors are scripted by hand-designed workflows rather than learned, and existing training methods optimize isolated components such as query rewriting or list ranking. As a result, agents rarely verify candidates, rarely use the user's history, and almost never recover from a rejected recommendation. We present , a reinforcement learning framework that trains the complete interaction policy of a tool-integrated recommendation agent. SimRec introduces two components. A task-anchored user simulator provides feedback on every recommendation: conditioned on the current task need rather than on the target item or a user profile, it states what the recommended item lacks without revealing the target, and it can be an external LLM or the recommender itself under self-play. A task-aligned reward optimizes the whole trajectory with GRPO, crediting natural retrieval quality and valid tool use alongside the final hit, with training-only target injection to handle sparse hits over large catalogs. On ESCI and Amazon C4, SimRec improves the NDCG@100 of a Qwen3-4B agent from and from , outperforming training-free agents on far larger backbones and task-aligned baselines. Analysis shows that the trained agent verifies candidates, uses preferences, and recovers from rejected recommendations far more often than its untrained counterpart. Our code is available at https://anonymous.4open.science/r/SimRec-6F0D/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.