COMPASS: Completion-based Process Advantage Supervision for Sequential Recommendation
Abstract
Delayed rewards make it difficult to determine which actions contributed to an outcome. Completing alternative actions from the same interaction prefix can supply useful comparisons, but the resulting returns reflect the policy used for completion. They need not remain suitable value targets as the learning agent improves. We propose COMPASS, which fits offline completion advantages and uses their centered predictions to supervise action-value differences during reinforcement learning. The agent continues to learn from environmental rewards through its TD objective, and the fitted model is reused across training runs. We analyze the resulting value updates and characterize how changes in continuation policy affect the validity of an action comparison. On sequential recommendation tasks, COMPASS improves DQN's mean held-out return by 20.7% and 43.1% on RL4RS-A and RL4RS-B. The same model improves IQN across eight seeds, and a slate-preference extension improves DDPG by 12.5% on KuaiSim. Experiments examine the choices involved in using this evidence, including pretraining, centering, and the duration of supervision. Deployment uses the learned policy alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.