Behavior Before Weights: Detecting Backdoors in Open-Source Adapters
Abstract
Adapters have become a popular medium for sharing specific tasks in open source communities due to their lightweight and plug-and-play nature. However, this sharing mechanism also introduces security risks, as malicious backdoors can be embedded in adapters and propagated through reuse. Existing detection methods often rely on anomalies in the parameter space and require training data or clean reference models, making them difficult to apply to untrusted third-party adapters. In this work, we study a detection setting without reference where the defender has access only to the adapter, its base model, and the downstream task description. We observe that backdoor activation can leave behavioral signatures of abnormal changes and accelerated convergence during inference. Based on this observation, we develop a detection method that explores the input space for tasks, identifies behavioral outliers, and further verifies their output consistency. Experiments on multiple downstream tasks, attack strategies, adapter variants, and ranks demonstrate the effectiveness of the proposed approach, while a deployment study on publicly available adapters further examines its practicality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.