On-Policy Representation Alignment in Preference Optimization
Abstract
Direct preference optimization (DPO) has become a simple and effective paradigm for aligning large language models (LLMs) with human preferences. However, DPO-style methods are fundamentally constrained by a mismatch between static offline preference data and the evolving behavior of the current policy. Recent works attempt to mitigate this mismatch by constructing preference supervision in the reward space or through data-side strategies. However, these constraints provide only a compressed view of the internal generation process and do not explicitly ensure that the current policy encodes sufficient preference information during generation. In this paper, we revisit this limitation from a representation perspective and propose On-Policy Representation Alignment (OPRA), which extends preference optimization from reward-level response ranking to representation-level preference alignment. Specifically, for an offline preference pair, we sample a short response prefix from the current policy as an anchor and encourage its response-level representation to be closer to that of the chosen response than to that of the rejected response. By jointly optimizing output-level preference consistency and on-policy representation alignment, OPRA provides a direct representation-level bridge between offline preference supervision and the policy's own autoregressive generation. Experiments on ten benchmarks demonstrate consistent improvements across preference-data settings and LLM backbones. Further analyses show that OPRA learns more preference-consistent representations and generalizes well across diverse settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.