Average Kernels are Better Kernels
Abstract
If multiple artists are asked to draw a circle by hand, each one will produce something slightly imperfect. Yet, the average of their sketches can look strikingly close to ideal. We investigate whether representations produced by separately trained models adhere to a similar principle. Precisely, we study the average model kernel: the Gram matrix computed from all model embeddings. Using metrics of representational similarity, i.e., Centered Kernel Alignment (CKA) and Mutual k-Nearest Neighbors (MKNN), we show that the average kernel increases in similarity with the kernel of a stronger model. We demonstrate this for vision models trained on different datasets, skewed datasets, models of different architectures, and even models trained on different tasks. Empirically, we even find that the similarity landscape with respect to teacher kernels is a convex basin. Finally, we show how averaging can be directly translated into improved model performance by using DiffKNN, a differentiable version of MKNN, to optimize a student network for representational similarity with the average kernel. We observe that the resulting student model increases in task accuracy with the number of models being averaged. Compared to approaches such as weight-averaging or multiple-distillation, which require that models start from the same checkpoint or share the same architecture, this technique is more general. More broadly, our findings hint at a more general framing that can be given for model-merging techniques, in which models can be thought to lie in the same loss basin with respect to their representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.