CompeteRL: Foresight–Hindsight Competitive Reinforcement Learning
Abstract
Experiential learning is becoming tightly integrated with agentic reinforcement learning (RL), which uses task outcomes to directly optimize experience utilization and extraction. However, existing methods separate experience extraction from utilization across rollouts, delaying feedback of newly extracted experience and leaving retrieved historical experience poorly adapted to the current task. To tackle the challenge, we introduce **Experience-in-the-Loop Learning**, a formulation that turns experience extraction and utilization into trainable decisions within the policy rollout. Building on this formulation, we propose CompeteRL, an online RL framework that integrates experience generation, utilization, and validation within a unified rollout, while improving experience quality through foresight–hindsight competition. Specifically, \ours adapts retrieved knowledge into foresight experience before acting and distills trajectory-grounded hindsight experience afterward. The two experiences compete through their guided execution returns, forming a closed loop of task-specific adaptation and immediate feedback. Extensive experiments across four agent benchmarks using Qwen2.5-7B and Qwen3-4B show that CompeteRL **(I) delivers strong task performance**, exceeding Skill1 by 6.1% on average with Qwen3-4B, and **(II) lies on the Pareto-optimal frontier** of cost and performance across four datasets, outperforming SkillRL by 7.4% on ALFWorld at comparable token cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.