acceptodds
Under review as a conference paper at ICLR 2027

DeepRecover: Learning Cross-Depth Safety Recovery for Large Language Models

Abstract

Large Language Models (LLMs) exhibit effective safety behavior at the start of generation, yet may fail to recover once adversarial prefixes push them into harmful trajectories. More importantly, such failures vary substantially with prefix depth, raising a key question: *Can safety recovery generalize beyond the generation depths observed during alignment?* To address this challenge, we propose `DeepRecover`, a privileged-anchor framework for depth-invariant safety recovery. `DeepRecover` constructs a privileged safe trajectory from the clean query and additional safety information, and uses it as a shared recovery anchor across prefix depths. Learning from this anchor faces two coupled mismatches: an *initial-state mismatch* between clean and prefix-corrupted contexts, and a *trajectory mismatch* as the attacked student evolves along a different recovery path from the privileged teacher. Accordingly, DeepRecover first uses transition tokens to redirect generation toward a recoverable state, and then applies unbalanced optimal transport to align subsequent recovery behavior with the privileged safety anchor flexibly. By coupling depth-specific recovery processes through the same safety reference, DeepRecover encourages a shared recovery capability beyond observed prefix depths. Experiments across four LLMs and multiple safety benchmarks demonstrate strong cross-depth generalization, with `DeepRecover` maintaining low attack success rates far beyond the observed depths. Code is available at [DeepRecover](https://anonymous.4open.science/r/DeepRecover-75D8).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.