How Much Evidence Is Needed to Trust Model Comparisons?
Abstract
Offline model comparisons depend on which candidates were allowed to compete. When historical membership is incomplete, a reconstructed list can turn an uncertain comparison into a seemingly precise winner. We study the evidence needed for a certified model choice. The key is shared uncertainty: membership changes can cancel between models. We prove that a comparison can require zero new records while fixing all selections requires linearly many. For single-query top- evaluation, we characterize the entire contrast through one shared boundary block. This yields an exact algorithm whose exponential dependence is local, and a linear-time regret formula for disjoint adjacent swaps. A matching search separation shows that exact computation and minimal evidence acquisition are distinct problems. For general model orders and multiple queries, joint disagreement and cutoff constraints provide efficient, monotone regret bounds. These certificates support comparison-guided acquisition and stopping before full recovery. Seven trained stock scorers exhibit 42 winner reversals across 132 pool-date comparisons; 23 incur regret exceeding 0.01. Frozen-score experiments repair all 23 above-tolerance errors while leaving more than half the candidate memberships unresolved on average. In a controlled width-two stock reranking task, the exact block certificate certifies 67/132 comparisons versus 5/132 for Cutoff and a block-span control at the same additional 25% record budget. Separate policy experiments save 53.1% of recommendation records at zero regret.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.