TrajEdit: Safety Editing with Multi-Anchor Trajectories
Abstract
Rapidly patching vulnerabilities in Large Language Models (LLMs) against evolving jailbreak attacks remains a critical challenge. While Knowledge Editing offers a resource-efficient alternative to retraining, existing methods indiscriminately update parameters over the entire response sequence, often disrupting general factual knowledge irrelevant to the safe direction associated with the target correction. This inefficiency leads to severe over-refusal on benign queries, catastrophic forgetting of general capabilities, and limited generalization to broader safety-related tasks. In this paper, we propose TrajEdit, a solver-agnostic framework that reframes safety correction as sparse, multi-anchor trajectory editing. It constructs a safe trajectory for each selected anchor and jointly edits the corresponding divergence points, thereby reducing updates to generic content. In our implementation, the first chunk is retained as an anchor, while perplexity-based uncertainty serves as a simple and effective option for selecting additional anchors. Extensive experiments across LLMs demonstrate that TrajEdit achieves superior defense success while maintaining a high non-refusal rate comparable to leading baselines, effectively reconciling the trade-off between safety and locality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.