acceptodds
Under review as a conference paper at ICLR 2027

Empirical and Certified Robustness Across Pre-training Regimes in Transfer Learning

Abstract

Transfer learning is a natural fit for safety-critical applications, where task-specific data is scarce but models have to be robust against adversarial perturbations. Robustness can be instilled during pre-training, during fine-tuning, or both, where robust pre-training is by far the most expensive of these. Yet it remains unclear when this expensive step is actually needed: prior work reaches opposing conclusions, and each study fixes a single adaptation mechanism while letting downstream data scale vary together with the target domain. In this paper, we take on a new comparative perspective to investigate empirical and certified robustness across pre-training regimes, fine-tuning objectives, adaptation mechanisms, and downstream data scales, under different adversary norms, also exposing trade-offs. Our findings show that the adaptation mechanism decides the answer, also reconciling prior work's settings and conclusions: With a frozen backbone, robustness can only be inherited from pre-training, while with an updateable backbone, robust fine-tuning instils robustness and standard fine-tuning largely erases it. We further show that robust pre-training nonetheless gains value for empirical robustness as downstream data becomes scarce. Finally, a layer-wise representational analysis reveals that comparable robustness can arise from distinct internal representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.