Robustness After Specialization: Installing Adversarial Robustness in Speech Foundation Model Descendants
Abstract
Speech foundation models are often specialized into task-specific descendants before adversarial robustness becomes a deployment requirement. We study *post-hoc robustness installation*: can robustness learned once at a foundation backbone be installed into an already fine-tuned descendant without downstream labels, the original fine-tuning pipeline, per-descendant adversarial training, or additional inference-time computation? From unlabeled source audio, we learn a task-agnostic backbone update . It reduces adversarial representation drift while explicitly minimizing clean-representation changes. For each descendant, we add the same update at a fixed, untuned scale. We then freeze the patched encoder and realign its linear readout on clean unlabeled audio. We evaluate three backbones on five speech tasks, producing 15 descendants through full encoder fine-tuning. With one shared update per backbone, seven-attack robust accuracy improves by 48.7 percentage points on average across the 12 classification descendants. The worst-task clean-accuracy loss is 7.3 percentage points. The method recovers 79.5% of the task-specific classification AT gain. Evaluation attacks the exact installed model, including with adaptive white-box attacks. Our analysis measures transferred representation stability and relates the remaining adversarial drift to downstream decision margins via an attack-relative sufficient criterion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.