acceptodds
Under review as a conference paper at ICLR 2027

Active Query Synthesis for Learning from Preferences

Abstract

Most response models for pairwise preference comparisons assume all feedback is equally reliable, overlooking that comparisons between near identical or highly dissimilar options are often ambiguous. We introduce a *confidence aware response model*, where response reliability depends on the similarity of the compared options. A user study we conducted shows it models human choices better than standard models. Separately, active learning reduces the cost of estimating preferences, but pool-based methods evaluate the available candidate queries each round. Building on our response model, we propose *Info-Synth*, an active query synthesis framework. It selects informative queries by maximizing mutual information in closed form, up to a one-dimensional search, avoiding pool-based search entirely. For fixed datasets, we introduce two approximation strategies, *Pair M-dist* and *Pair Opt-dist*, to select the closest available pairs of points. We evaluate our framework on synthetic preference learning, a constrained text-summary dataset, and continuous-space controller gain tuning for a simulated robot. In each case, it matches or exceeds the sample efficiency of pool-based active learning while reducing the computational cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.