acceptodds
Under review as a conference paper at ICLR 2027

Decision-Focused Acquisition of Preference Data for Learning Inference-Time Verifiers

Abstract

Modern Large Language models increasingly use their compute budget for inference-time computation. In applications such as code generation and tool-use, this often involves generating multiple candidate responses and learning a verifier that applies a decision rule such as best-of-K selection, reranking or shortlisting, or escalating to an expert to decide the best response. In such cases, the final utility of the response depends both on the verifier’s score and the decision operator applied. In this paper, we consider the problem of active learning to train such verifiers. Existing active learning methods for preference data select the comparisons that most reduce uncertainty in the verifier’s parameters, and are indifferent to which directions of parameter space the deployed decision actually consults. Instead, we propose the Decision Focused Acquisition (DFA) criterion, a c-optimal experimental design that targets the decision. We assume access to a small Decision Set where utilities are known and the DFA criterion selects the comparisons that most reduce the posterior variance of a linearised decision loss. This formulation applies to any decision operator that has a differentiable surrogate and we instantiate it for best-of-K selection, top-m shortlisting, and accept-or-escalate cascades. Across four datasets spanning preference reranking, code generation, and tool-use agents, we show that our method outperforms existing active preference learning methods, yielding verifiers with higher decision utility under fixed acquisition budget. Lastly, we show that the benefit of DFA is largest at small and medium budgets, where acquisition is most critical, and that it comes from acquiring comparisons at the decision boundary of the deployed operator.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.