Pareto-GRPO: Prompt-Local Pareto Credit Assignment for Efficient Search-Augmented Reasoning
Abstract
Group-relative reinforcement learning works well when a verifier supplies a clear scalar reward, but search agents expose a missing interface: trajectories can reach the same answer while using very different amounts of search. Fixed scalarizations resolve such ties by imposing one global exchange rate between correctness and cost, even though the useful amount of search is prompt dependent. We introduce Pareto-GRPO, which keeps the critic-free GRPO update but orders each prompt's rollout group by Pareto dominance over correctness and search efficiency. A bounded rank utility then refines the group-normalized correctness advantage. This ordinal construction is invariant to coordinate-wise monotone reward transforms and admits an exact first-order objective characterization. In a controlled Qwen2.5-7B comparison with identical initialization, data, retrieval, optimization, and validation-only model selection, Pareto-GRPO reaches 0.465 average exact match with 1.78 searches and 0.286 hypervolume. The strongest tuned scalar baseline reaches 0.441, 1.85, and 0.258, respectively; all five seed-wise accuracy differences favor Pareto-GRPO. The advantage persists under enforced search caps, alternative cost definitions, and a three-objective extension. Across seven benchmarks and two model scales, Pareto credit assignment improves accuracy, efficiency, and run-to-run stability together.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.