Revisiting sharpness and Model Generalization through Local Rademacher Complexity
Abstract
Sharpness-aware minimization (SAM) and its variants can improve test accuracy by minimizing the maximum increase in empirical loss under local model-parameter perturbations. This improvement is commonly attributed to the view that higher loss sharpness is strongly associated with poorer generalization. However, several studies challenge this view and report weak or even opposite correlations. Regardless of the correlation, one more fundamental question is whether the sharpness alone can explain model generalization. Using local Rademacher complexity, we develop a new theorem on excess risk, which shows that, in comparison with reducing the sharpness alone, jointly reducing sharpness and complexity will lead to a compacter excess-risk bound. Focusing on Rademacher complexity within the theorem, we derive a bound which increases as the sample-wise classification margins decrease. Furthermore, we build a Rademacher-guided plug-in mechanism, termed Rad, that adjusts the direction of parameter perturbations according to sample-wise classification margins. Extensive experiments on multiple datasets and model architectures demonstrate that Rad combined with SAM or its variants improves the test accuracy and reduces the magnitude of the train–test accuracy gap. Grounded in local Rademacher complexity, our work offers a new interpretation of the relationship between loss sharpness and model generalization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.