Differentiable Kernel Methods with Error Certificates for Deep Learning Pipelines
Abstract
Modern deep learning pipelines deliver strong accuracy but little per-prediction error control, whereas classical kernel methods come with a mature approximation theory that does not, by default, scale to or plug into such pipelines. We close that gap with two constructs. Sparse kernels make kernel ridge regression scalable: a localized lazy variant that defers computation to inference and solves only small local systems. Differentiable kernel layers make it modular: the same regressor exposed as an end-to-end trainable module, in a doubly-projected form that admits distinct embeddings for the query and the centers while keeping a symmetric positive-definite kernel, and whose three parameter sets (representations, targets, evaluation points) may each be fixed or learned. The two compose, and the resulting design space is broader than end-to-end training alone: training-free transfer, nonlinear probing, hybrid and purely kernel models, classification and policy heads. Our thesis is that these constructs let a deep model carry a computable, geometry-aware error indicator, evaluated at each query and propagated through the pipeline, a kernel-specialized estimate that conformal prediction can build on and a concrete step towards auditable predictions. Across vision, reinforcement learning, out-of-distribution detection and tabular data, such components integrate at competitive cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.