acceptodds
Under review as a conference paper at ICLR 2027

CO-OIPO: Coverage-Optimized Online IPO for Process Rewards under Intransitive Preferences

Abstract

Process reward models (PRMs) provide signals for evaluating and guiding intermediate steps in multi-step reasoning. Existing PRMs typically adopt a Bradley–Terry (BT) model trained on preference pairs, with the training objective of distinguishing the better from the worse within a pair of samples. However, in multi-hop knowledge graph question answering, multiple paths may correspond to the correct answer, leaving some intermediate steps without a clear distinction between better and worse. Furthermore, multiple preference pairs may form long-chain preference relations, thereby causing a sample to appear in multiple positions within the preference relations, which gives rise to contradictions. To address the above problem, we propose formulating preference learning as a game over reasoning candidates, with a regularized Nash equilibrium as the learning objective. This formulation accommodates a broader class of pairwise preference relations, allowing two samples without clear preference to receive similar exploration probabilities. Furthermore, to make more effective use of a limited comparison budget, we propose Coverage-Optimized Online Identity Preference Optimization (CO-OIPO), which optimizes pair allocation within a sampled candidate pool. We prove that, with sufficient sampling coverage, an expressive enough policy model, and full optimization in each round, CO-OIPO converges at a fixed point corresponding to a regularized Nash equilibrium of the preference game.Experiments show that CO-OIPO achieves the highest performance on most metrics; the controlled experiments further demonstrate more effective utilization of samples without clear preference, and correctness of theoretical analysis.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.