Latent Intent, Bounded Evidence: Reinforcement Learning for Language-Model Recommenders
Abstract
We introduce PILOT, a method that trains a language-model recommender with re- inforcement learning while keeping a language-model critic from inventing reward that the interaction logs cannot support. The difficulty is structural: long-horizon recommendation quality is exactly the quantity that offline logs report least reli- ably, so any critic accurate enough to be useful is also free enough to be gamed, and policies that chase it drift into slates that score well and serve users poorly. PILOT resolves this by making the critic’s licence explicit. A frozen language model infers a discrete latent intent for the current session; a semantic-identifier policy decodes a slate conditioned on that intent; and the reward is composed by allow- ing a language-model judge to adjust a logged reward ensemble only within the confidence interval that the ensemble’s own epistemic uncertainty leaves open. We prove that this composition has a bias envelope proportional to the reward model’s uncertainty, and that the resulting policy admits a safe-improvement guarantee against any comparator, whereas an unbounded judge leaves an irreducible value gap no amount of data removes. Across three public recommendation corpora and a calibrated session simulator, PILOT improves both next-item accuracy and long-horizon session utility over sequential, language-model, and reinforcement- learning baselines, and—unlike the baselines it is compared against—its true ses- sion utility does not collapse as the policy is pushed further from its reference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.