AdaPPO: Resolving Dynamics Asymmetry in LLM Alignment via Adaptive Gating
Abstract
Proximal policy optimization (PPO) is widely used for aligning large language models, but it typically updates the actor and critic at every step despite their different optimization timescales. We observe and analyze a timescale asymmetry: under a local Polyak–Łojasiewicz condition, the critic converges exponentially, whereas the actor admits a sublinear stochastic convergence bound. This suggests that synchronous updates can incur redundant critic computation once the critic enters a low-gradient regime, while freezing it for too long can make its value estimates stale and degrade actor optimization. We therefore bound critic error accumulation under bounded value drift, which motivates a staleness-aware freeze schedule. Based on this analysis, AdaPPO adaptively refreshes the critic using three signals—performance, KL divergence, and entropy—evaluated jointly, with a maximum-freeze safeguard. We further introduce AdaPPO*, which augments the critic with privileged information. Experiments on four RLVR benchmarks show that AdaPPO reduces FLOPs by up to 25% and reaches target performance up to faster in training steps. Our code is available at https://anonymous.4open.science/r/AdaPPO-E375.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.