acceptodds
Under review as a conference paper at ICLR 2027

Contrastive Policy Optimization: Optimizing von Neumann Entropy Across Queries for Unsupervised Foundation Model Alignment

Abstract

Reinforcement learning (RL) post-training has emerged as a dominant paradigm for unlocking complex reasoning and continuous generative capabilities in foundation models. However, standard post-training methods (such as PPO, GRPO, and DPO) rely heavily on expensive supervised signals, including ground-truth reasoning trajectories, verified labels, or calibrated reward models. While recent unsupervised heuristics, such as Entropy Minimized Policy Optimization for language reasoning and Atomic Policy Optimization for 3D atomic generation, attempt to eliminate external supervision, they evaluate generations strictly in isolation within individual queries. Consequently, these isolated intra-query methods suffer from severe vulnerabilities: , where the policy converges to degenerate, uninformative generic templates or universal crystal packing motifs regardless of input conditions, and susceptibility to induced by ad-hoc cluster approximation errors. To address these foundational bottlenecks, we propose Contrastive Policy Optimization (CPO), a unified, theoretically grounded framework for fully unsupervised foundation model alignment. Instead of optimizing isolated Shannon entropy per prompt, CPO models the joint generation distribution across a batch of diverse queries as a global quantum state represented by a normalized density operator . By leveraging contrastive learning to optimize the , CPO simultaneously drives two synergistic objectives: (1) , forcing the policy to form decisive, high-confidence consensus modes on each specific query, and (2) , enforcing maximal semantic distinction and diversity across different queries. We derive an exact, analytical policy gradient and advantage formulation that eliminates the need for heuristic clustering, external reward models, or supervised coordinates. Extensive experiments across both continuous geometric generation and discrete language reasoning demonstrate the decisive superiority of CPO. Our theoretical and empirical results demonstrate that cross-query von Neumann entropy optimization provides a universal, mathematically rigorous foundation for unsupervised foundation model alignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.