acceptodds
Under review as a conference paper at ICLR 2027

Energy-based DAG Structure Refinement for Large Language Model Preference Alignment

Abstract

Aligning large language models (LLMs) with human preferences remains a central challenge in modern AI. Reward models (RMs) serve as proxies for human judgment but are highly susceptible to overoptimization and reward hacking. RM ensembles provide complementary evaluations, but scalar pooling compresses their disagreements and graph aggregation can retain conflicting edges. In this paper, we propose **E**nergy-based **D**AG s**T**ructure ref**I**nement (EDIT) to learn a prompt-specific preference graph from multiple RMs. A structural equation model (SEM) supplies a tractable graph-conditioned density, and a residual energy-based model (EBM) fits joint patterns beyond this base model. We train the structural module and energy function with SEM likelihood and noise contrastive estimation (NCE) using adaptive negatives from a conditional flow. Acyclicity regularization and edge pruning yield a directed acyclic graph (DAG) for consistent ranking. Extensive experiments on RMB and AlpacaFarm demonstrate state-of-the-art performance in both response selection and policy optimization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.