acceptodds
Under review as a conference paper at ICLR 2027

Reliable Characterization of Local Model Complexity with Valid Localization

Abstract

Understanding training dynamics and learned models requires more than evaluating loss or predictive performance. Modern deep neural networks(DNNs) are often overparameterized, and solutions with similar performance can have different local structure, which motivates a principled measure of model complexity around a solution. Singular learning theory provides such a framework, and the local learning coefficient(LLC) quantifies local model complexity, including cases where parameter symmetries and redundancies make distinct parameter values represent the same predictive function. A scalable sampling-based approach estimates LLC around a target solution using Gaussian localization. Because the localization penalty is part of the sampled target, its strength can change the measured local model complexity. Weak localization can admit remote low-loss regions, whereas excessive localization can impose an artificial contraction scale and distort the target complexity. By resolving the loss and localizer in a common geometric representation, we derive model-agnostic localization conditions under which localization-based measurements asymptotically converge to the LLC of the target solution across a broad class of analytic models, without estimating model-specific singular geometry. We further establish classwise counterexamples showing that neither the lower nor upper localization requirement admits a uniform relaxation over this analytic model class. The guarantee is asymptotic. For finite computations, we separate the discrepancy between the population localized measurement and the target LLC from empirical and numerical errors, and derive exact diagnostics that quantify sensitivity to localization. Experiments on truth-known singular models and DNN training trajectories show that theory-guided localization yields reliable LLC measurements and reveals reproducible complexity changes during grokking and learning-rate transitions. These results establish when localization-based LLC measurements can be interpreted as properties of a target solution rather than consequences of the localization choice.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.