Complementary PAVE Specifications for Characterizing and Identifying LLM Capabilities
Abstract
Parameter Vector (PAVE) *specifications* encode each candidate large language model (LLM) and each task as fine-tuning updates of a common auxiliary model, enabling LLM selection without evaluating candidates on the target task. Extending PAVE beyond models derived from a common base, we find that auxiliary models provide complementary views of model–task compatibility: the same behavioral observations can yield accurate model comparisons in one parameter space but misleading comparisons in another. To connect these views to actual model performance, we prove that, for a fixed auxiliary model and under explicit assumptions, the probability that two tasks agree on a model pair's true performance ordering is non-decreasing in their nonnegative PAVE cosine similarity, up to a quantified approximation error. This result motivates Reference-Guided Multi-Auxiliary Voting (RG-MAV), which uses evaluated reference tasks that are most informative for the target in each auxiliary model's parameter space to weight and combine model comparisons separately for each candidate pair. Extensive experiments on heterogeneous LLMs and target tasks from six domains support the benefit of reference locality and show robustness to specification-training seeds and individual auxiliary-model removal. RG-MAV achieves an average Kendall's of , which exceeds the of a fixed best auxiliary model selected on reference tasks and the of EmbedLLM, and reaches over 96% of the attained by a post-hoc task-wise oracle. These results demonstrate that RG-MAV can effectively select and combine accurate views, and that reusable parameter-space characterizations, guided by reference-task evidence, can support effective model selection before any candidate is run on the target task.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.