Efficient Policy-Space Regression via Reparameterized Optimal Targets
Abstract
Supervised ranking-based methods such as Direct Preference Optimization (DPO) have shown that bypassing complex RL dynamics can improve the stability and efficiency of post-training. However, these methods impose supervision primarily in the reward space, and their guarantees do not generally ensure alignment of the model policy with the optimal policy in policy space, potentially introducing bias. To address this issue, we propose Minimum Cross-entropy Policy Optimization (MCPO), which performs regression directly in the policy space. MCPO first reparameterizes the optimal policy into a more tractable target, and then regresses the policy toward it using a low-variance cross-entropy-style objective. Theoretically, we show that the optimal policy remains a stationary point of the MCPO objective, preserving the desired optimum under practical variance control. Moreover, under idealized tabular softmax models, MCPO guarantees monotonic progress toward the optimal policy when clipping is non-degenerate; without clipping, it further yields the closest single finite-step update to the optimal policy among all objectives. Empirically, MCPO consistently outperforms all baselines across all evaluated settings. In preference alignment, it achieves better performance using only 50% of the training steps, establishing an efficient paradigm for LLM post-training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.