When Uniform Tie-Breaking Preserves Global Score Structure
Abstract
Pairwise preference optimization learns from binary comparison outcomes, but many annotations also include ties. Prior work has incorporated ties into preference objectives and characterized their effect on population learning targets. Uniform tie-breaking preserves every comparison, but the conditions under which the resulting probabilities admit a single score vector have not been characterized. We establish this compatibility boundary for Davidson comparisons with a shared tie strength through cycle closure, which requires pairwise log-odds to sum to zero around every comparison cycle. We derive the exact residual sign and tight fixed-span extrema, and use inverse-link closure to characterize all feasible tie strengths, including multiple isolated solutions. The analysis explains DPO’s population projection and separates reward-scale changes that temperature calibration can correct from distortions in score direction. On controlled graphs and shared policies, calibration reduces source-policy regret, and cycle mismatch remains distinct from the additional fitting cost of parameter sharing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.