acceptodds
Under review as a conference paper at ICLR 2027

Nash Credit Routing for Optimizing Responses Beyond GRPO

Abstract

A GRPO update can improve its sampled objective while reducing the likelihood of an individual successful response. Repeated, directionally similar successes can also dominate the positive update through multiplicity alone. We introduce Nash Credit Routing (NCR), a constrained correction to GRPO that reallocates a fixed positive-credit budget across successful responses and selectively refunds penalties on failed-response segments. The correction remains near GRPO, preserves at least its first-order sampled-objective improvement, and causes no additional first-order likelihood harm to any sampled success. NCR assigns positive bargaining powers by minimizing directional collision under a multiplicity-calibrated ratio constraint, then maximizes weighted Nash welfare over normalized likelihood gains. Across the evaluated models and mathematics benchmarks, NCR achieves higher macro-average pass@1 and pass@128 than the compared baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.