Interact, Remember, and Act: Preference and Progress-Driven Reinforcement Learning for LLM Agents
Abstract
Effective collaboration between users and large language model (LLM) agents depends on both successful task execution and interactions that accommodate user preferences. However, existing agentic reinforcement learning (RL) approaches incorporating user simulator provide limited modeling of agents' capabilities and user experience, with no unified environment across heterogeneous tasks, and efforts to improve user experience and long-horizon credit assignment remain largely separate. In order to address these issues, we introduce PPD-RL, a multi-turn agentic RL framework jointly driven by user preferences and task progress. Specifically, PPD-RL defines a standardized action space for agents encompassing clarification queries, tool use, memory revisions, and final answers, together with unified interaction environments for training and inference. Moreover, to enable personalized assistance across sessions, the agent maintains an explicit user profile, updates it when certain attributes in the profile are missing or drifting, and conditions subsequent interactions on the corrected memory. Concurrently, to improve these behaviors, PPD-RL combines the task goal, rule-based memory feedback, and an LLM-judged preference score as RL rewards. It then distributes credit across multi-turns using positive rewards and measurable task progress, linking the learning signal to actions that advance task completion. We train and evaluate PPD-RL across various gym families spanning travel planning, customer service, and collaborative programming, measuring both task performance and user experience. Experimental results demonstrate that PPD-RL outperforms existing methods in this area and the larger models included in our evaluation on both dimensions. Code is available at https://anonymous.4open.science/r/PPD-RL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.