Partial Policy Gradients for Conversational RL in LLMs
Abstract
Learning to act sequentially in an unknown environment can be formulated using reinforcement learning and solved by policy gradients. We propose a natural approach for trading off the bias and variance in policy gradients. The key idea is to optimize for a subset of future rewards: smaller subsets represent simpler policies, which can be learned more efficiently because their empirical gradient estimates have lower variance. Our approach encompasses various policy classes, such as full planning, greedy, and -step lookahead policies. We evaluate the policies empirically in the offline setting on conversational tasks of persona alignment and agentic tool-calling, and directly measure gradient variance across lookahead horizons. Different policies excel in different problems, reflecting their different characteristics and highlighting the importance of our studied trade-off.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.