Disentangling Safety and Capability in Large Language Models via Causal Interventional Decomposition
Abstract
Safety alignment in large language models can deteriorate during downstream post-training because task learning and safety control remain coupled through shared Transformer computation. Existing safety-preserving methods primarily reduce geometric overlap between task-related and safety-related parameter directions, but geometric separation does not directly indicate whether the resulting functions are truly independent. To address this issue, we propose LCID, Linear-to-Causal Interventional Disentanglement, which decomposes post-training updates into shared, task-private, and safety-private components and uses controlled branch interventions to quantify their target and cross-functional effects. LCID directly suppresses task-to-safety and safety-to-utility cross-effects while preserving the intended contributions of each private branch and anchoring the complete model to its original aligned safety behavior. Experiments on Qwen3-4B, Llama-3.1-8B, and GLM-4-9B show that LCID substantially improves the safety-utility trade-off and consistently reduces unsafe responses across multiple safety benchmarks. Prompt-level paired analysis further demonstrates that these safety improvements are systematic across harmful inputs. Intervention analysis shows that near-orthogonal parameter updates can still exhibit substantial functional cross-talk, whereas LCID achieves markedly lower cross-effects while maintaining strong target effects. These results establish functional intervention as a direct criterion for disentangling safety and capability during post-training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.