acceptodds
Under review as a conference paper at ICLR 2027

TrajEdit: Trajectory-Level Model Editing for Persistent Jailbreak Backdoors

Abstract

Jailbreak backdoors can selectively subvert the refusal behavior of safety-aligned large language models (LLMs) under predefined trigger conditions. Existing editing-based attacks improve backdoor activation by manipulating trigger-related representations or behavioral objectives. However, existing approaches primarily operate on localized states or fixed behavioral targets, without explicitly modeling how the induced behavior evolves during autoregressive generation. As decoding proceeds, the induced behavior may gradually weaken, making it difficult to sustain over the full response. We introduce TrajEdit, a trajectory-level model editing framework for persistent jailbreak backdoors. TrajEdit combines three complementary designs: (1) *trajectory targeting*, which derives decoding-stage-aware behavioral targets from paired target-refusal trajectories; (2) *persistent retrieval*, which preserves edit accessibility throughout decoding; and (3)*rollout adaptation*, which stabilizes the target behavior under post-edit generation dynamics. Experiments across multiple safety-aligned LLMs and harmful-request benchmarks show that TrajEdit improves attack effectiveness and long-horizon behavioral persistence while maintaining trigger selectivity and preserving general model capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.