Finite-Sample Guarantees for Convex Neural Preference Learning
Abstract
Pairwise preference data are commonly used to train reward models. Yet existing statistical analyses of pairwise preference learning often assume linear reward heads, leaving finite-sample guarantees for nonlinear heads less developed. We study a Bradley–Terry estimator that fits a two-layer ReLU reward head on fixed representations through a convex activation-pattern formulation. The estimator jointly fits the preference likelihood and selects among retained activation patterns using a block-sparsity penalty. Conditioning on the representations and retained patterns, we analyze squared error in reward differences on the observed comparisons. For rewards supported on at most of patterns in dimension with samples, we establish a minimax lower bound of order under a sparse design condition. We further derive a prediction error bound of order for the convex estimator under standard conditions. When is known, a sharper calibration of the same estimator achieves the minimax lower bound rate under a stronger restricted eigenvalue condition. Experiments with synthetic rewards, language-model representations, and five preference datasets assess reward recovery and held-out preference prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.