Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models
Abstract
Large language models (LLMs) may require additional defenses after deployment as risks and governance requirements evolve. Subsequent defenses can interact with earlier defenses, raising the question of their sequential compatibility. We study this question with ConflictEval, evaluating 144 ordered compositions of six defenses spanning safety, privacy, and fairness across six models. The resulting interactions are heterogeneous across defense pairs, application orders, and models. Notably, some compositions exhibit defense conflicts: the subsequent defense improves its target objective while weakening protection established by the earlier defense. We investigate these interactions through the directional compatibility of defense-induced changes in risk-relevant representations. Across 11 selected cases, we use activation interventions to assess which defense-induced representational changes support protection and examine how subsequent defenses affect these changes. We find that, in some conflict cases, subsequent defenses counteract representational changes supporting earlier protection, providing evidence for one possible pathway to defense conflicts. Building on this analysis, we propose Conflict-Triggered Directional Retention (CTDR), which penalizes opposing shifts along directions supporting earlier protection. On six selected conflicting compositions from this analysis, CTDR reduces first-objective regression while maintaining positive gains on the subsequent defense objective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.