DACPO: Drift-Adaptive Constrained Policy Optimization for Cue-Free Safety Recovery in Large Reasoning Models
Abstract
Large reasoning models (LRMs) can reproduce safety corrections from supervised traces, yet this does not ensure that they will recognize and reverse an unfamiliar adversarial shift in their own reasoning. We call the discrepancy between cue-conditioned correction competence and cue-free self-correction the correction activation gap. We introduce Drift-Adaptive Constrained Policy Optimization (DACPO), a two-phase framework that combines supervised fine-tuning (SFT) and reinforcement learning (RL) for robust recovery from adversarial reasoning drift. The first phase uses SFT to establish a safety correction prior from adversarial correction trajectories. The second phase removes target corrections and uses RL to optimize cue-free recovery from policy-dependent worst-tail drift states. DACPO couples conditional value-at-risk state selection with trajectory-level recovery rewards and a persistent archive of newly discovered and historical failures, while an over-refusal constraint prevents safety gains from collapsing into indiscriminate refusal. Experiments across multiple LRMs and benchmarks show improved recovery in both directions while preserving direct safety and general reasoning utility. Together, these findings identify autonomous correction activation, rather than correction imitation alone, as a central requirement for robust reasoning alignment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.