acceptodds
Under review as a conference paper at ICLR 2027

PEARL: Process-Experience Aware Reinforcement Learning for Multi-Turn Agents

Abstract

Reinforcement learning for multi-turn agents often relies on sparse outcome signals that indicate whether a trajectory succeeds, but provide limited guidance about which intermediate decisions should be corrected. This becomes especially problematic for failed trajectories, which may still contain many valid actions and therefore should not be treated as uniformly negative experience. We propose PEARL (Process-Experience Aware Reinforcement Learning), a framework that learns from recurring failure patterns across past interactions. Rather than penalizing an entire failed trajectory, PEARL identifies failure patterns that repeatedly appear in unsuccessful behavior and uses successful experience to remove patterns that are not uniquely associated with failure. The remaining failure-specific patterns are then used to provide more targeted negative supervision during reinforcement learning. On AppWorld, PEARL+ARPO consistently improves over ARPO on both Normal and Challenge tasks at the 4B and 8B scales. On ALFWorld, PEARL achieves 97.9% mean success at 4B, outperforming the strongest baseline by 2.7 points, while remaining competitive at 8B. Extensive ablations and process analyses further validate the effectiveness of PEARL's selective failure-memory design.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.