acceptodds
Under review as a conference paper at ICLR 2027

Partial Policy Gradients for Conversational RL in LLMs

Abstract

Learning to act sequentially in an unknown environment can be formulated using reinforcement learning and solved by policy gradients. We propose a natural approach for trading off the bias and variance in policy gradients. The key idea is to optimize for a subset of future rewards: smaller subsets represent simpler policies, which can be learned more efficiently because their empirical gradient estimates have lower variance. Our approach encompasses various policy classes, such as full planning, greedy, and -step lookahead policies. We evaluate the policies empirically in the offline setting on conversational tasks of persona alignment and agentic tool-calling, and directly measure gradient variance across lookahead horizons. Different policies excel in different problems, reflecting their different characteristics and highlighting the importance of our studied trade-off.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.