Learned Shared Space for Kernel-based Meta-learning
Abstract
Kernel methods are attractive for few-shot regression because squared-loss predictors can be computed in closed form from a small set of observations. Their performance depends on how well the kernel captures inherent structure across tasks, motivating the learning of a kernel representation through meta-learning. However, learning the kernel alone does not provide a learned initial function as in standard gradient-based meta-learning. We develop a kernel meta-learning method, based on inducing points, to learn both a common function space and initial function within it. The learned initial function is then transferred to each new task and adapted by functional gradient descent. We characterize the resulting finite adaptation trajectory to better understand the role of gradient steps during adaptation. We further derive excess transfer risk bounds that quantify the effects of approximation, initialization, and adaptation steps. Under the stated assumptions including boundedness, regularity and spectral decay condition of the kernels, we provide an early-stopping rule and the resulting risk bounds, and describe the trade-off between the approximation error of the learned common space along with the cost of learning the inducing points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.