acceptodds
Under review as a conference paper at ICLR 2027

How Much of a Language Model Group's Collective Failure Is Visible in Its Pairs?

Abstract

How much do the pairwise failure rates of the models in a language-model group reveal about them failing together? Individual and pairwise rates, which form a pairwise table, do not in general identify the collective failure rate, the share of items on which every model fails. However, across 48 public evaluation pools from five releases, completions of the pairwise table approximate the collective failure rate closely. For groups of eight models, the range of collective failure rates allowed by the pairwise table is on average 16 percentage points wide, more than twice a group's observed rate, so the table alone pins that rate down only to within ±120% of its value. Yet the least-structured completion, pairwise maximum entropy, reconstructs that rate with 15.9% mean relative error (1.1 percentage points), and a Gaussian copula gives the same qualitative result. Where pools have enough items to measure it, most of the dependence among model failures is pairwise. The regularity has limits: error rises with group size and on held-out items when collective failures are scarce. One use follows. Across the 48 pools, groups chosen by maximum entropy from a pairwise table fail together less often on held-out items than groups of the most accurate models, on average and in most pools (3.3 percentage points at eight models), also when the choice is restricted to the strongest candidates. Exploratory studies on SWE-bench Verified and RewardBench 2 are consistent with this. The results hold within an evaluation, not across evaluations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.