acceptodds
Under review as a conference paper at ICLR 2027

Robustness After Specialization: Installing Adversarial Robustness in Speech Foundation Model Descendants

Abstract

Speech foundation models are often specialized into task-specific descendants before adversarial robustness becomes a deployment requirement. We study *post-hoc robustness installation*: can robustness learned once at a foundation backbone be installed into an already fine-tuned descendant without downstream labels, the original fine-tuning pipeline, per-descendant adversarial training, or additional inference-time computation? From unlabeled source audio, we learn a task-agnostic backbone update . It reduces adversarial representation drift while explicitly minimizing clean-representation changes. For each descendant, we add the same update at a fixed, untuned scale. We then freeze the patched encoder and realign its linear readout on clean unlabeled audio. We evaluate three backbones on five speech tasks, producing 15 descendants through full encoder fine-tuning. With one shared update per backbone, seven-attack robust accuracy improves by 48.7 percentage points on average across the 12 classification descendants. The worst-task clean-accuracy loss is 7.3 percentage points. The method recovers 79.5% of the task-specific classification AT gain. Evaluation attacks the exact installed model, including with adaptive white-box attacks. Our analysis measures transferred representation stability and relates the remaining adversarial drift to downstream decision margins via an attack-relative sufficient criterion.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.