acceptodds
Under review as a conference paper at ICLR 2027

Which Geometry Predicts Generalization Depends on What You Vary — Activation or Optimizer

Abstract

Recent work introduced R\'enyi sharpness, a reparameterization-invariant measure of loss-Hessian geometry (Zhang et al., 2026), and reported that it correlates strongly and consistently with the generalization gap, the difference between training and test accuracy. That evidence, however, comes from varying the optimizer while the activation function is held fixed. We invert this setup. We fix the optimizer to stochastic gradient descent (SGD), with a transfer check under a sharpness-aware optimizer, and vary the activation function across seven common functions. Across multiple datasets and architectures, we compare weight-space, input-space, and activation-level measures by their signed Kendall rank correlation with the activation-induced generalization gap. In this setting, R\'enyi sharpness is weak and changes sign across architectures and datasets (mean Kendall of in our main comparison). In contrast, the input Jacobian, which measures how sensitive the output is to the input, and a mechanistically motivated activation-level quantity are positively associated with the gap in every setting of our main comparison (mean of each). This quantity, the negative-branch leak, is the mean squared activation slope on negative pre-activations, , where is the activation and the pre-activation. A controlled intervention that gates only the negative branch provides causal evidence that the leak contributes to input sensitivity. Guided by this mechanism, we control the leak in two ways. First, we reduce it directly by scaling down the activation's negative branch, with no added loss term. This improves clean accuracy for every activation in our main setting, by up to points, and by up to points on a wide residual network (WRN-28-10). Second, we add a leak penalty to the training loss, together with the established input-Jacobian penalty (Hoffman et al., 2019). The leak term mainly affects clean accuracy, while the input-Jacobian term drives corruption robustness. Which geometry predicts generalization depends on the axis along which models vary.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.