acceptodds
Under review as a conference paper at ICLR 2027

When Priors Veil the Threat: Backdoor Defense via Channel Activation Inspection for Pretrained Federated Learning

Abstract

Existing defenses against backdoor attacks in federated learning aim to identify malicious clients and typically rely on the discrepancies among model gradient updates to distinguish them from benign ones. In this paper, we empirically show that when the global model is initialized with pretrained backbones, the updates similarity between benign and malicious clients can be substantially amplified, thereby rendering existing defenses ineffective. To bridge this gap, we investigate client model discrepancies from the perspective of channel activations and uncover a key observation: benign models generally exhibit consistent channel activation magnitudes across samples whereas backdoored models show abnormal activation variations. Motivated by this, we propose CAS (Channel Abnormality Score), a new malicious detection metric quantifying the abnormality within the channel activation pattern of a client model via the notion of term frequency-inverse document frequency (TF-IDF). Specifically, CAS treats channel activations as words and the feature map of each class as a document. When a client contains channels that contribute abnormally to documents as measured by CAS, the server identifies the client model as backdoored and excludes it from global aggregation. Extensive experiments show that CAS significantly outperforms existing defenses under the pretrained federated learning setting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.