From Dilution to Alignment: Understanding and Mitigating Long-Context Safety Degradation
Abstract
As large language models (LLMs) become more capable, their safety risks become increasingly consequential. Existing safety research has focused primarily on short-context settings, leaving long-context safety alignment largely underexplored. Prior studies show that model safety degrades markedly as inputs grow longer, even when the surrounding context is entirely benign. We use safety-vector projection to trace this degradation through the attention and MLP pathways. Our analysis identifies two coupled mechanisms: input-side risk dilution, which reduces attention to risky queries, and downstream safety-perception decline in MLP layers. Guided by these findings, we propose SALSA (Safety-Aware Long-context Self-distillation Alignment). Its first component, Safety-Aware Group Sequence Policy Optimization (SA-GSPO), strengthens safety perception through response-level safety judgment. Its second component, Segment Summary On-Policy Self-Distillation (SS-OPSD), re-encodes diluted risk in segment summaries generated within reasoning and distills this behavior into a prompt-free policy. An exponential moving average (EMA) anchor on benign samples helps preserve response rates and general capabilities. Extensive experiments show that SALSA substantially improves long-context safety over Plain GSPO baselines and mitigates length-induced safety degradation, bringing long- and short-context safety rates into close agreement. General capability and response rates remain largely unaffected in both long- and short-context settings. Further mechanistic analysis also confirms that SALSA alleviates the attention- and MLP-level failures underlying long-context safety degradation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.