acceptodds
Under review as a conference paper at ICLR 2027

When Uniform Tie-Breaking Preserves Global Score Structure

Abstract

Pairwise preference optimization learns from binary comparison outcomes, but many annotations also include ties. Prior work has incorporated ties into preference objectives and characterized their effect on population learning targets. Uniform tie-breaking preserves every comparison, but the conditions under which the resulting probabilities admit a single score vector have not been characterized. We establish this compatibility boundary for Davidson comparisons with a shared tie strength through cycle closure, which requires pairwise log-odds to sum to zero around every comparison cycle. We derive the exact residual sign and tight fixed-span extrema, and use inverse-link closure to characterize all feasible tie strengths, including multiple isolated solutions. The analysis explains DPO’s population projection and separates reward-scale changes that temperature calibration can correct from distortions in score direction. On controlled graphs and shared policies, calibration reduces source-policy regret, and cycle mismatch remains distinct from the additional fitting cost of parameter sharing.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.