Precision Thresholds for Shared Evaluation Caches
Abstract
Model evaluation is often staged: a library is scored once on a shared set of examples, and later reports ask for more precise comparisons without rebuilding the entire evaluation table. We ask how far such a finite cache can be reused before fresh evaluation must scale with the size of the library. We consider complete i.i.d. score vectors in , arbitrary unknown cross-model dependence, and a fixed collection of mean differences that must satisfy a uniform marginal MSE target under a hard cap on fresh work. The key obstruction is residual candidate uncertainty: additional measurements of a shared reference improve every contrast, but cannot uniformly remove the uncertainty left in each candidate's finite cache. For a unit-price reference star with candidates, this yields an exact threshold separating bounded, , and linear fresh work in ; matching lower bounds hold for fully adaptive acquisition, with critical work and a transition window. A stronger layout-wise accuracy contract changes the critical law to , with an integer crossover governed by . For partial observations on general comparison graphs, we give an exact bounded-score reduction to edgewise union counts and asymptotically exact high-precision work limits.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.