Anchor Reconstruction for Efficient Model-Task Evaluation: Identifiability and Stability
Abstract
As the numbers of models and tasks grow rapidly, dense evaluation becomes increasingly costly since each new model or task may require many additional evaluations. Anchor reconstruction reduces this burden by using approximately low-rank model-task structure to estimate the full performance vector of a new task or model from evaluations on a few selected models or tasks, called anchors. This setting raises three questions: what determines reconstruction across nonunique factorizations, when do sparse anchor scores uniquely determine the full vector, and how do errors arise, propagate, and become amplified? For a fixed rank- approximation, we show that the performance subspace rather than a particular basis is the intrinsic object. Anchor reconstruction is identifiable exactly when coordinate restriction is injective on this subspace, while stability is governed by the anchor minimum gain. We derive an end-to-end reconstruction error bound separating target-subspace mismatch, historical-matrix perturbation, and anchor-evaluation noise, with inverse gain amplifying all three. This explains why approximate low rank and enough anchors alone do not ensure stable reconstruction. Experiments on a matrix and a public leaderboard illustrate the rank-budget boundary for identifiability and support the predicted roles of error pathways and inverse-gain amplification in reconstruction stability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.