acceptodds
Under review as a conference paper at ICLR 2027

Preference Optimization for Distribution-Level Objectives

Abstract

Post-training typically targets individual-level objectives, which evaluate a policy through the expected value of a per-response reward. Broader goals like diversity, calibration, and fairness, however, demand distribution-level objectives that depend on the collective outputs and probabilities of a policy. In this work, we consider jointly optimizing individual- and distribution-level objectives, and make the following contributions. (i) We formulate the joint optimization problem and derive a per-response linearization reward for the distribution-level objective. This motivates the development of an iterative surrogate optimization method with dynamic linearization, which we show provably converges to a solution of the joint objective. (ii) We adapt existing post-training methods to the surrogate problems, introducing a general framework for incorporating distribution-level objectives into standard post-training pipelines. Experiments across three testbeds demonstrate improved distribution-level outcomes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.