acceptodds
Under review as a conference paper at ICLR 2027

Efficient Policy-Space Regression via Reparameterized Optimal Targets

Abstract

Supervised ranking-based methods such as Direct Preference Optimization (DPO) have shown that bypassing complex RL dynamics can improve the stability and efficiency of post-training. However, these methods impose supervision primarily in the reward space, and their guarantees do not generally ensure alignment of the model policy with the optimal policy in policy space, potentially introducing bias. To address this issue, we propose Minimum Cross-entropy Policy Optimization (MCPO), which performs regression directly in the policy space. MCPO first reparameterizes the optimal policy into a more tractable target, and then regresses the policy toward it using a low-variance cross-entropy-style objective. Theoretically, we show that the optimal policy remains a stationary point of the MCPO objective, preserving the desired optimum under practical variance control. Moreover, under idealized tabular softmax models, MCPO guarantees monotonic progress toward the optimal policy when clipping is non-degenerate; without clipping, it further yields the closest single finite-step update to the optimal policy among all objectives. Empirically, MCPO consistently outperforms all baselines across all evaluated settings. In preference alignment, it achieves better performance using only 50% of the training steps, establishing an efficient paradigm for LLM post-training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.