acceptodds
Under review as a conference paper at ICLR 2027

Which AI Agent Actually Wins? Active Query Generation for High-Confidence Selection

Abstract

Choosing an artificial intelligence (AI) agent for deployment can depend on a small subset of queries that distinguish otherwise similar candidates. These diagnostic queries may be rare and unknown in advance. Evaluating agents on fixed benchmarks or randomly sampled queries can therefore spend much of the evaluation budget on cases that do little to distinguish them. To address this problem, we formulate a framework for selecting the best agent under a limited evaluation budget, by adaptively choosing which queries to evaluate. Within this framework, we propose Diagnostic Agent Selection via Exploration (DASE), which directs evaluation toward distinguishing the most competitive candidates. At each round, DASE identifies the current leader and a challenger, then generates queries likely to distinguish them using previous evaluations. The generator can be adapted to the application, for example by using a large language model conditioned on the evaluation history. DASE selects one of these queries, evaluates all agents on it, and updates its estimates. Theoretically, we provide lower bounds on the number of diagnostic queries selected and the probability of selecting the best agent. Empirically, we evaluate DASE in three settings: simulated agent selection, mathematical reasoning, and clinical decision-making. Across these settings, DASE generally achieves higher best-agent selection accuracy than the tested baselines, particularly at larger evaluation budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.