acceptodds
Under review as a conference paper at ICLR 2027

Defending Multilingual Refusal Capability Against Three-Layer Cascading Attacks with Behavior-Invariant LoRA Distillation

Abstract

Safety-aligned language models often encode refusal behavior in internal activation directions, but these representations can be fragile across layers, prompt positions, and languages. We introduce Three-Layer Cascading Attack (TLCA), a representation-guided attack that produces ordinary textual prefixes and suffixes. Once constructed, the attack requires neither access to model internals nor any modification of model parameters, activations, or decoding procedures. We then propose Behavior-Invariant LoRA Distillation (BILD), which freezes the base model and trains low-rank adapters to reproduce its clean behavior under adversarial perturbations. We show that a fixed-template LoRA adversarial-training baseline can suppress harmful compliance under attack but also causes severe benign-utility degradation and repetitive refusal behavior. In contrast, BILD distills the model's own clean refusals and helpful answers, thereby avoiding universal-refusal collapse while preserving instruction-specific response diversity in multilingual contexts. We additionally compare BILD with Refusal Feature Adversarial Training (ReFAT), a representation-level adversarial-training method, under the same TLCA evaluation protocol. Extensive multilingual experiments show that TLCA substantially weakens refusal behavior across model families, whereas BILD consistently restores robustness while maintaining strong compliance on harmless requests.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.