Tailoring Transferability Estimation Evaluation to Application Preferences
Abstract
As the number of available pretrained models continues to grow, identifying the most suitable source models for a downstream task becomes increasingly difficult. Transferability estimation (TE) methods address this problem by ranking pretrained models without any exhaustive fine-tuning. Existing works evaluate these methods by measuring their agreement with a ground-truth ranking induced by a single score, mostly fine-tuned accuracy in classification. However, the accuracy may not be aligned with the objective of the downstream application, which may assign different importance to particular classes or errors, and thereby induce a different ranking of the same models. In this work, we argue that the evaluation of TE methods should explicitly account for downstream application preferences. To this end, we introduce a new framework in which ground-truth rankings are induced by performance scores tailored to the downstream application preferences. These scores form a family that includes the accuracy as a special case. When application preferences are uncertain, we evaluate TE methods across a distribution of application-conditioned ground-truth rankings. Through experiments, we show that conclusions drawn from conventional TE benchmarks do not necessarily hold across downstream preferences: ground-truth model rankings, their agreement with TE methods, as well as the induced ranking of TE methods vary with the downstream preferences. For binary classification, we further show how this variability can be visualized over the entire space of downstream preferences. The code is available in supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.