acceptodds
Under review as a conference paper at ICLR 2027

LLM Benchmark Compression via Maximum Independent Set Prompt Selection on Proximity Graphs

Abstract

Evaluating LLMs across comprehensive benchmarks is expensive, motivating benchmark compression: replacing a full benchmark with a compact subset that is cheaper to run. Whether such a method is usable comes down to fidelity—when does the subset reproduce the full-benchmark LLM ranking, and when does it diverge? We model each benchmark as a prompt proximity graph—nodes are prompts, and an edge joins two prompts whose embeddings are close enough to be treated as semantically redundant—and apply Maximum Independent Set (MIS) solvers to select a maximally diverse, non-redundant subset in which no semantic region is doubly counted. We evaluated four MIS solvers across six embeddings, three distances, six percentile thresholds, and four benchmarks (GPQA, IFEval, MMLU-Pro, Omni-MATH) covering 66 LLMs, spanning 2,558 configurations. Two axes emerge and move almost independently. First, selection is consistent: different random seeds yield near-identical rankings (Kendall's in 84.5% of stochastic configurations; mean ), so the procedure itself is reliable. Second, agreement with the full-benchmark ranking is governed by how aggressively the graph is pruned, giving a controllable efficiency–fidelity trade-off: mild thresholds (–) preserve the full ranking (–) while removing 64–82% of prompts, whereas aggressive pruning (–) trades fidelity for compression (–). Because the axes separate, the grid becomes a recipe: the threshold decides whether a subset stays reliable-and-faithful, and the embedding decides how it fails once it diverges. Divergence concentrates at high thresholds and on the more discretized benchmarks (GPQA, IFEval); when it occurs, a category-level analysis shows MIS redistributes prompt mass toward uniform coverage, which we report descriptively. We recommend ReduMIS with qwen8b embeddings, standardized-Euclidean distance, and a threshold as a strong default operating point.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.