Scaling Laws and Mitigation of Memorization in Kolmogorov Arnold Networks
Abstract
Malign memorization limits the deployment of Kolmogorov-Arnold Networks (KANs) under corrupted labels, particularly in scientific machine learning where observational measurements contain label corruptions. While empirical studies observe this performance degradation, existing work does not explain how the defining KAN controls, grid size and spline degree , govern the resulting population damage. We analyze this behavior by decomposing memorization into corruption fitting and its population footprint , giving . In the localized regime of a two-layer B-spline KAN, grid refinement accelerates corruption fitting while contracting its footprint, whereas increasing degree slows fitting and broadens the footprint, yielding the matched-fit scaling . This decomposition explains why fixed-budget training often report worse memorization at larger , even though matched-fit memorization becomes strictly more localized. Guided by this analysis, we propose Curvature-Guided Basis Perturbations (CURB), which provably slows corruption fitting by perturbing curvature-sensitive basis directions during training. Across function regression, scientific operator learning (Darcy flow), and classification benchmarks, CURB reduces label corruption fitting, achieves competitive or superior accuracy to SAM and ASAM at lower per-update cost, and transfers successfully across diverse KAN basis families. In addition, CURB consistently leads trained models to flatter empirical loss landscapes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.