DAPO: Dual Advantage Policy Optimization for Aligning Large Language Models
Abstract
Reinforcement learning from human feedback (RLHF) has emerged as a cornerstone technique for aligning large language models (LLMs) with human preferences. However, existing approaches such as Proximal Policy Optimization (PPO) suffer from two fundamental limitations: reward hacking—where the model exploits gaps in the reward model rather than learning genuinely preferred behaviors—and conflation of distinct objectives, particularly the trade-off between linguistic fluency and semantic alignment. In this paper, we propose Dual Advantage Policy Optimization (DAPO), a novel RLHF training framework that disentangles these objectives by maintaining two separate advantage streams: a Fluency Advantage (AˆF t ), computed via a frozen language model scoring perplexity, and an Alignment Advantage (AˆA t ), derived from a trained reward model. We introduce an adaptive gating network gϕ that dynamically weighs these two streams at each decoding step, conditioned on the generation context. Extensive experiments on Anthropic HH-RLHF, AlpacaEval, MT-Bench, and TruthfulQA demonstrate that DAPO achieves state-of-the-art performance across all benchmarks, outperforming PPO by +7.1% win-rate and DPO by +8.4% win-rate on AlpacaEval, while maintaining significantly lower KL divergence from the reference policy (0.81 vs. 1.42). Ablation studies confirm the contribution of each component, and qualitative analysis reveals that DAPO generates responses that are simultaneously more fluent and more aligned with human values.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.