acceptodds
Under review as a conference paper at ICLR 2027

Benchmark Prediction with Sparse and Unpaired Cached Responses

Abstract

Evaluating a generative model on all queries from all available benchmark tasks is prohibitively expensive. As such, models are often evaluated on different subsets of tasks and/or different subsets of queries from a given task. The resulting collection of responses and the corresponding (model, task) score matrix are sparse and unpaired. In this paper, we introduce the Product Kernel Perspective Space (\PKPS), a collection of low-dimensional representations of models based on their embedded responses. \PKPS generalizes existing methods for representing black-box generative models by comparing responses to similar – rather than just identical – queries and is therefore able to utilize sparse and unpaired information. Theoretically, we show that benchmark prediction using \PKPS is query-efficient relative to methods that only use information from dense and paired portions of the evaluation data. Empirically, we show that \PKPS outperforms baselines on benchmark prediction across two large benchmark suites.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.