acceptodds
Under review as a conference paper at ICLR 2027

PLUME: PERSISTENT LEARNING FROM UNIFIED MULTI-ROLLOUT EXPERIENCE FOR AGENTIC REIN- FORCEMENT LEARNING

Abstract

Large language models are increasingly trained to act in multi-step agent environments, and group-relative policy optimization (GRPO) with verifiable rewards is widely used to train such agents. However, its trajectory-level advantages provide limited guidance on which decisions contributed to the outcomes, and the sampled trajectories are discarded after policy updates without their textual experience being distilled into reusable lessons. We introduce PLUME (Persistent Learning from Unified Multi-Rollout Experience), a framework that learns from rollout groups not only in numerical space through GRPO’s relative advantages but also in text space through contrastive analysis. Within each task, PLUME compares alternative trajectories to extract lessons from successes and failures. It then compares these task-specific lessons across a batch to distill generalizable skills, which are accumulated in a persistent skill bank across training steps. Both trajectory analysis and skill generalization are performed by the policy model itself, requiring neither a stronger external teacher nor externally supplied skills. The accumulated skills guide policy optimization through self-distillation, without being inserted into the policy’s rollout or inference-time prompts. PLUME thus yields two artifacts: an improved policy and a standalone skill bank for transfer. Experiments on ALFWorld, WebShop, and search-augmented QA show consistent gains across three model scales. The resulting skill bank also transfers zero-shot to larger models from different families, improving performance in almost all evaluations across the three benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.