A NOVEL BACKDOOR ATTACK AND DEFENSE TAR- GETING COST-EFFECTIVENESS OF LARGE LANGUAGE MODELS
Abstract
While LLM capabilities for diverse fields, such as computer programming, mathematical and scientific discovery, and the law, continue to grow, of considerable concern is the cost — not only of building them, and of the infrastructure (capital expenditures for datacenters) for huge frontier models, but also of using them (operational expenditures). A substantial component of this cost is the token length of the model's response to a given input prompt. In this work, we propose a novel backdoor attack against instruction-following in LLMs specifically designed to result in overly long LLM responses, i.e., this attack is at the intersection of LLM security and its economy. Moreover, we show that this attack is quite effective even when the backdoor trigger phrase is chosen to contain a synonym of "Be concise" — such an attack is readily inadvertently triggered. Finally, we develop a post-training backdoor defense, based on the recent detection paradigm "class subspace orthogonalization", both to detect this attack and to invert the ground-truth trigger phrase. Experiments on text summarization domains show the potency both of the attack and of the defense we mount against it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.