APAF: Adaptive Pareto Advantage Fusion for Multi-Reward Policy Optimization
Abstract
As large language models (LLMs) become increasingly capable, post-training reinforcement learning (RL) often relies on multi-dimensional reward signals to shape model behavior. However, fusing potentially conflicting rewards into a single update signal remains a central challenge. Common reward-fusion approaches either linearly combine reward dimensions into a scalar, allowing gains in one dimension to be offset by losses in another, or filter out conflicted rollouts, discarding supervision that can still provide useful learning signal. We propose Adaptive Pareto Advantage Fusion (APAF), a reward-fusion algorithm for multi-reward policy optimization. Within each rollout group, APAF uses Pareto dominance to capture the relative structure among samples while retaining a magnitude-based advantage that preserves reward magnitude information. Its core is a relative-margin control mechanism: for each query, APAF estimates the strength of pairwise Pareto evidence and adaptively sets a bounded mixing weight, determining how much structural signal to trust. In this way, APAF exploits reliable group-level Pareto relations without discarding rollouts or suppressing reward dimensions. Experiments on tool use and safety alignment show that APAF consistently outperforms representative reward-fusion baselines, demonstrating that Pareto-aware advantage fusion provides a principled, general, and effective mechanism for multi-objective post-training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.