Learning Reachable Plans: PlanGap-Guided Distillation and Plan-Contrastive Policy Optimization for Reasoning
Abstract
Reinforcement learning with verifiable rewards improves long chain-of-thought (Long-CoT) reasoning in large language models (LLMs). However, outcome-supervised Group Relative Policy Optimization (GRPO) derives a single response-level advantage from each rollout's terminal reward and applies it across all tokens. To support fine-grinded advantages, we introduce a Plan-and-Action structure that places a concise high-level plan before detailed execution. We first validate this structure through a PlanGap-guided distillation, which extracts plans from verified teacher solutions and selects those with smaller teacher–student gaps in mean plan negative log-likelihood to prioritize strategies accessible to the student. We then propose Plan-Contrastive Group Relative Policy Optimization (PC-GRPO), which operates independently of the distillation stage and requires no SFT cold start. Under a fixed verifier budget, PC-GRPO samples multiple plans and multiple executions per plan, using their outcomes to estimate plan reliability and construct separate learning signals for plan selection and plan-conditional execution. It further contrasts plans with higher and lower estimated values using only plan tokens, providing plan-specific supervision without privileged solutions, an external teacher, or additional full-trajectory scoring for the contrastive signal. Experiments across diverse reasoning domains, model families, and scales show improved reasoning performance and more efficient plan supervision and policy optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.