acceptodds
Under review as a conference paper at ICLR 2027

Certificate Tightness Is a Training Property

Abstract

A deterministic robustness certificate rests on a global Lipschitz bound, usually the product of the layers’ spectral norms. On a standard network under standard training, that bound exceeds the sensitivity the network exhibits by orders of magnitude. We study the ratio of the two, the gap. The bound certifies the true constant from above, and any exhibited Jacobian certifies it from below. So the gap is the factor between the two ends of a certified enclosure. No verifier returning a global constant can improve the certified radius by more than that factor. This paper characterizes what sets the gap. It does not build a certified system. Its one intervention leaves the architecture and the loss fixed. We track the gap along training over 380 runs on four datasets. At a fixed convolutional architecture, training sets it. At learning rates fixed in advance, the gap grows 3× under SGD, 60× under AdamW and 2,300× under Muon. Tuned rates compress that spread to one order, SGD still mildest and Muon below AdamW. At those fixed rates on ImageNet-64, the spread widens to ten orders. Every quantitative prediction was registered before its test run, with two marked exceptions, and every miss is printed. The gap factors exactly into misalignment between layers and slack from units the activations switch off. Projecting each layer onto a spectral-norm cap after every step removes the slack term. Under Muon it moves the gap from 3.2 million to 7.4, at 0.55 clean accuracy against 0.88. The product bound is then within a factor of about seven of the best radius any global-constant verifier could grant. Repair after training recovers a factor of two to six, against five orders. Under the cap the optimizers’ roles invert. At the operating budget and the rates fixed in advance, SGD and AdamW collapse toward rank one and stop learning. Muon, which inflates the gap most when free, keeps training. We name this the constraint inversion. The cap moves only the bound, and verified accuracy also needs a trained margin. On four published 1-Lipschitz constructions at one fixed bound, Muon verifies more than SGD, with disjoint seed ranges. Training decides how much of a tight bound is usable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.