acceptodds
Under review as a conference paper at ICLR 2027

Extreme Region Policy Distillation

Abstract

Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: on-policy methods require repeatedly generating expensive trajectories, while off-policy reuse introduces distribution mismatch. We find that aggressive optimization on fixed data brings rapid initial gains. Continued updates then shift token probabilities and reduce entropy after performance plateaus. Stronger KL regularization reduces both divergence and attainable performance. This motivates Extreme Region Policy Distillation (ERPD), a two-stage framework that separates aggressive rollout reuse from the final policy update. Stage 1 performs weakly constrained optimization on a fixed rollout batch to construct a strong policy-difference signal. Stage 2 restarts from the policy that generated the batch and optimizes the fixed teacher–reference log-probability difference with a clipped update, retaining useful policy changes while limiting divergence. The resulting policy matches or exceeds its teacher with substantially smaller KL divergence. ERPD also accommodates strong and weak teachers: when the Stage 1 teacher remains below the base model, contrasts between intermediate and recovered checkpoints still provide effective supervision. Across mathematical reasoning benchmarks, iterative ERPD raises Qwen3-4B-Thinking-2507 accuracy on HMMT Nov 25 from 66.6% to 79.0%; at larger scale, it raises Qwen3.5-27B exact-match accuracy on APEX2025 from 2.08% to 14.6%, while improving the other reported transfer benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.