BiProxy: Bilevel Optimization for LLM Backdoor Persistence via Proxy Adaptation
Abstract
Open-weight LLMs are commonly fine-tuned after release, which can weaken backdoors implanted before release. Achieving persistent backdoors is particularly challenging because an upstream attacker generally has no access to the downstream data, task, or adaptation procedure. In this paper, we investigate whether persistence learned through a lightweight adaptation proxy can transfer across unseen downstream updates. To this end, we propose BiProxy, a bilevel framework that separates a persistent implant LoRA from a disposable proxy LoRA. During each training episode, the proxy undergoes benign adaptation while the implant remains fixed, after which the implant is optimized to preserve trigger-conditioned behavior in the proxy-adapted state while maintaining clean behavior. The proxy is reset between episodes and discarded before release, and a first-order update avoids differentiating through its adaptation trajectory. Across two primary LLMs, two attack objectives, and three downstream tasks, BiProxy achieves post-adaptation attack success rates of 98.2%–100% under full-parameter fine-tuning and 93.2%–100% under LoRA fine-tuning in the default settings, while generally preserving comparable benign utility. Removing proxy adaptation substantially reduces persistence in some settings despite similar release-time attack success. These results show that a low-rank proxy can provide transferable training signals for persistent backdoors without reproducing the downstream fine-tuning trajectory. Our code is available in an anonymized repository: https://anonymous.4open.science/r/ICLR_2027Code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.