PODA: Partial-Order Advantage Estimation for Multi-Reward LLM Alignment
Abstract
Multi-reward reinforcement learning for large language models usually compresses every reward vector into a scalar advantage with a fixed linear rule. This is convenient, but it inserts pairwise preferences even when two rollouts trade off different objectives. We introduce PODA, a prompt-local advantage estimator that instead constructs a dominance graph, combines an inverse dominance-count score with an antisymmetric dominance margin, and leaves the PPO/GRPO policy objective unchanged. PODA is a credit-assignment rule applied after rewards and any task priorities have been specified; it does not infer user preferences or learn a Pareto coverage set. Across mathematical reasoning and tool calling with two and three rewards, and across Qwen2.5, Qwen3, and Qwen3.5 models, PODA improves both primary performance and constraint satisfaction over equally tuned linear, lexicographic, rank-sum, additive-indicator, hypervolume, GRPO, GDPO, and DAPO comparators. On the representative two-reward settings, it improves the strongest matched indicator baseline by 0.91 math-accuracy points and 0.55 tool-accuracy points; the corresponding three-reward gains are 0.9 and 0.8 points. Paired three-seed confidence intervals are positive in the representative settings, while the estimator adds about 1% step time at group size 16. Training and dominance diagnostics show the largest gains when reward conflict and graph sparsity make fixed scalarization or count-only scoring least informative.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.