Contextual Credit Policy Optimization
Abstract
Group-based reinforcement learning enables large language model agents to learn from sparse outcome rewards in long-horizon tasks. Recent methods derive step-level credit by comparing observation-matched experiences across trajectories. However, identical observations can occur in different contexts, making some peers more informative for a given decision. Moreover, favorable terminal outcomes can obscure redundant actions or detours within successful trajectories. We propose Contextual Credit Policy Optimization (CCPO), which combines contextual peer weighting with history-conditioned return credit and realized future progress. CCPO combines frozen LLM representations and accumulated trajectory statistics to weight observation-matched peers from other trajectories by contextual similarity. Historical credit compares the realized return against the weighted peer baseline, while future credit measures the change in contextual outcome value over a short realized continuation. CCPO reuses collected rollouts without an auxiliary value critic or additional environment interactions; future information is used only during training. With Qwen2.5-1.5B-Instruct and Qwen2.5-7B-Instruct, CCPO reaches up to 97.4 ± 0.5% success on ALFWorld and 79.4 ± 1.3% on WebShop; across variants, 7B gains over GRPO are 12.3–15.4 and 5.5–6.5 percentage points. Ablations support query-trajectory exclusion, combined historical and future credit, and structured context representation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.