acceptodds
Under review as a conference paper at ICLR 2027

Graded Equivalence Certificates for Preference Optimisation

Abstract

Preference optimisation has dozens of objectives published as distinct methods, and several frameworks that place them in one parameterised family. Family membership does not tell a practitioner whether switching methods changes anything. Every objective we study gives each response a score, takes the gap between two scores as a margin, and applies a link, a one-dimensional loss, to the margin. We ask at which level two objectives are the same and separate four. Two objectives can share the ranking they induce over responses, the score up to a per-prompt scale and shift, the loss, or the policy that training converges to. Each level has a certificate, an algebraic check that two objectives agree, and a witness, a few numbers that show they do not. The first three levels form a strict chain. The fourth follows from a shared loss and from nothing weaker. The optimum is set by the loss's target map, the score gap the loss drives a labelled pair towards as a function of how often annotators prefer one response. Methods that share DPO's score have different target maps and converge to different policies. Applying the certificates to the methods in current use shows that a shared loss is rare and a shared ranking is common. On a trained model the certified relationships hold exactly, and four links that share DPO's score order their margins as their target maps predict. We also measure the gap the certificates leave open, between scores that use a reference model and scores that use the policy alone. A closed-form experiment shows that when the labelled pairs form a chain, the choice of link alone can reverse which response the policy prefers, and that comparing every response against one anchor removes the effect for every link.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.