Uncovering the Limits of IBP-based Certified Training through Scaling
Abstract
While deep neural networks have enabled remarkable advances across a wide range of supervised learning tasks, their susceptibility to adversarial attacks remains a major concern. Certified training methods based on interval bound propagation (IBP) produce models with mathematically rigorous robustness guarantees, but typically suffer from degraded performance on clean data. While recent methods improved this trade-off by training on both adversarial examples and IBP bounds, progress has stalled, suggesting that IBP-based certified training may be approaching its limits. Here, we show that the progress that has been achieved came at the cost of severe overfitting, which traditional IBP training suffered less from. Thus, to uncover new performance limits of these methods, for the first time, we employ large-scale synthetic datasets and conduct the first systematic scaling study of IBP-based certified training across data and compute. We demonstrate that scaling yields performance improvements comparable to several years of algorithmic progress; for instance, on CIFAR-10 with , we obtain a model with 83.65% clean and 65.90% certified accuracy, substantially surpassing the previous state of the art of 79.87% and 64.54%, respectively. Analysing the resulting models, we find that these improvements are accompanied by changes in the distribution of unstable ReLUs and bound propagation tightness. At the same time, we reveal more fundamental limits of IBP-based certified training: performance quickly saturates with scaling, the gained robustness conflicts with propagation tightness, and stronger models incur increasingly prohibitive verification costs, making verification systems a bottleneck for certifying robust models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.