acceptodds
Under review as a conference paper at ICLR 2027

FIAPBA: Federated Instruction-Assisted Persistent Backdoor Attacks against Large Language Models

Abstract

Federated instruction tuning provides a practical way to adapt large language models (LLMs) without centralizing private client data. However, it also creates a vulnerable training surface for malicious clients. Existing backdoor and instruction attacks are mainly designed for centralized training or single-client poisoning. In federated LLM tuning, such attacks can degrade local utility, become easier to detect, and rapidly lose effectiveness after subsequent clean updates. We propose FIAPBA, a Federated Instruction-Assisted Persistent Backdoor Attack against LLMs under federated learning (FL). FIAPBA decomposes poisoning into two heterogeneous channels, an instance-level backdoor channel and an instruction-level attack channel. These two channels are deployed on different compromised clients and are coupled through federated aggregation, enabling the global model to learn the malicious target behavior while avoiding abnormal performance degradation on any single client. To improve persistence, FIAPBA dynamically identifies layers with relatively small parameter changes during backdoor tuning and embeds malicious updates into these more stable layers, making the attack less likely to be overwritten by later clean training. Extensive experiments on classification and generation datasets, with TinyLlama, Alpaca, and LLaMA2, show that FIAPBA achieves near-perfect attack success in many settings while preserving clean accuracy. Compared with representative backdoor and federated poisoning baselines, FIAPBA provides a stronger balance among attack effectiveness, stealthiness, and persistence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.