Model Replacement under Limited Adaptation
Abstract
When a model is retired, unavailable, or too costly to call, one may still want to reproduce its behavior by querying other available models and retaining a small amount of target-specific information that specifies how their responses should be combined. Which available models provide the best replacement when this retained information is limited? We show that the answer can depend fundamentally on the budget. Even when the original model is known completely, two sets of available models can achieve exactly the same replacement error with unrestricted adaptation yet reverse their optimal ordering when the retained state increases from one to two bits, with arbitrary encoders and nonlinear predictors allowed. Replacement error separates into an irreducible component and the error of compressing the target behavior predictable from training observations and the available models’ responses. Under Gaussian predictable behavior, the covariance spectrum of this behavior determines the entire information–error tradeoff and gives a necessary and sufficient condition for one model set to dominate another at every budget. Ex- periments with language models on three multiple-choice tasks show that matching retained-state budgets changes both adapter comparisons and which models are selected, while compact states can preserve much of the attainable performance: on Gemma/HellaSwag, 848 target-specific bits achieve test total variation 0.1951, compared with 0.1933 for full-precision kernel ridge regression.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.