acceptodds
Under review as a conference paper at ICLR 2027

Mechanism-Preserving LoRA: Causal Internal Constraints for Capability-Preserving Detoxification

Abstract

Large Language Models remain prone to generating toxic content, motivating detoxification methods that improve safety without compromising model utility. Parameter-efficient fine-tuning with LoRA is attractive for this setting due to its low cost and zero-overhead deployment, but existing approaches optimize primarily for behavioral outcomes: cross-entropy specifies what the model should output while imposing no constraint on how its internal representations change. Consequently, optimization can alter network mechanisms beyond what is required for detoxification, causing substantial perplexity inflation and downstream capability loss. We propose Mechanism-Preserving LoRA (MP-LoRA), a training framework that constrains not only the desired output but also the internal computation used to produce it. MP-LoRA first discovers and causally validates a toxicity direction, then records the layer-wise activation shifts induced by this intervention as a reference trajectory. A LoRA adapter is trained to reproduce this trajectory through auxiliary direction and magnitude-matching losses alongside cross-entropy, transferring a validated inference-time mechanism into the model parameters. Across five LLMs spanning 7B–32B parameters and two languages, MP-LoRA achieves up to 86% toxicity reduction while retaining 97–100% of accuracy on the seven zero-shot benchmarks evaluated. Ablations show that direction matching is important for capability preservation, while magnitude matching provides additional detoxification strength.The method further generalizes to instruction-tuned models and politeness transfer using the same pipeline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.