acceptodds
Under review as a conference paper at ICLR 2027

From Pairwise Evidence to Set Utility: A Controlled Meta-Evaluation of Homogeneity Measures for LLM Response Selection

Abstract

Selecting a small subset of large language model responses requires removing redundancy while preserving distinct valid alternatives. However, strong pairwise homogeneity prediction for the intended use does not guarantee good subset selection; even with perfect pairwise rankings, the selection pipeline can still incur avoidable coverage loss under a fixed selection budget. To explain this gap, we introduce a controlled meta-evaluation framework that holds responses fixed while varying their intended use across four task families with verifiable answers. We trace how decision thresholds fixed on development data interact with clustering and budgeted selection to shape final coverage. When evaluation pools different uses, a purpose-only control can achieve high AUC without inspecting response content, yet cannot distinguish alternatives within the target use. Input-order enumeration further shows that incorrect merges or splits reduce coverage only under certain repetition patterns, orderings, and budgets. On Natural-50, a collection of open-ended responses with human-annotated functional units, we evaluate coverage when retaining three of six responses per query. Transferred to 30 test queries, the pipeline selected by pooled AUC on controlled tasks incurs a mean per-query coverage shortfall of 8.36 percentage points relative to the optimum attainable from the same candidates under the same budget (95% CI: 1.94–16.03), with losses concentrated in six queries.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.