acceptodds
Under review as a conference paper at ICLR 2027

Self-Activating LLM Backdoors

Abstract

LLM Backdoors pose a critical and growing security risk to LLM systems. However, existing attacks, where adversaries choose specific triggers to activate malicious behaviour, assume either an interaction of the adversary with the deployed model or require a correct prediction of triggers present at deployment. Either assumption significantly decreases the practicality of current backdoors. We overcome these limitations by constructing self-activating backdoors, where the model itself generates an accumulative trigger signal. Specifically, we train LLMs to generate and recognize their own watermarked outputs and use watermark strength as a trigger signal. This threat model is especially relevant for coding agents, where models repeatedly modify and re-ingest their own generated code. Agents can thus gain user trust by behaving seemingly normal for extended periods before autonomously entering a malicious state, e.g., by introducing a malicious dependency or deleting files. Self-activating backdoors can also be combined with other trigger signals to increase attack specificity and reduce detection likelihood. Across 3 model families and 3 scenarios, our method achieves near-100% attack success rate while preserving benign behavior prior to activation. We demonstrate this effect in both controlled experiments and end-to-end coding tasks executed using a realistic agent harness. Our results expand the typical LLM backdoor threat model by showing that adversaries neither need externally supplied nor anticipated triggers, calling for a more realistic evaluation of LLM backdoors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.