acceptodds
Under review as a conference paper at ICLR 2027

Greedy alignment methods are probably (and provably) more efficient than you think

Abstract

AI alignment methods based on Bradley-Terry reward-model fitting and KL-regularized Alignment methods are remarkably effective in both offline and online preference-learning pipelines, yet existing online and offline rate guarantees can seem pessimistic relative to their empirical performance. We argue that this mismatch reflects a distinction between two learning targets: KL-regularized regret measures recovery of the entire finite-temperature policy, whereas reward learning and best-response identification ask whether the learned reward induces the correct top-ranked response. To isolate this reward-learning component, we study the traditional temperature-zero regret criterion, which evaluates only the top-ranked response at inference time. Under this decision-centric notion of performance, we prove that exact reference-logged offline RLHF achieves exponentially small expected temperature-zero regret, and therefore reaches expected regret at most \(\epsilon\) with \(O(\log(1/\epsilon))\) offline comparisons. We also prove that standard greedy online alignment methods, including online RLHF and online DPO, achieve constant \((O(1))\) cumulative temperature-zero regret. By separating selector identification from full soft-policy recovery, our results provide a sharper theoretical explanation for the empirical efficiency of greedy alignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.