SAE-Locked LoRA: Edit-Specific Sparse-Feature Gradient Routing for Low-Interference Knowledge Editing
Abstract
Knowledge edits should change a requested behavior without perturbing unrelated behavior, yet a low-rank adapter can alter broad internal directions despite updating few parameters. We introduce SAE-Locked Adaptive LoRA, which uses a frozen sparse autoencoder (SAE) as an edit-specific coordinate system for LoRA optimization. Normalized edit-loss gradients select the smallest bounded latent set covering a target attribution mass; only the edit gradient is routed through this fixed set, while preservation gradients use a separate unmasked pass. On CounterFact, the method retains comparable success to vanilla LoRA (99.6% versus 99.8%), improves locality (89.5% versus 63.7%), and reduces KL drift (0.16 versus 0.65). Under an objective-matched comparison, it gains 7.5 locality points over PCA (95% CI [6.7, 8.3]). It reaches 49.8% exact match and 28.5% micro on matched ReCoE and retains 91.8% success after 100 sequential edits, versus 18.2% for MEMIT and 8.4% for vanilla LoRA. Cross-backbone and bypass-path tests support the routing mechanism while delimiting its architectural scope.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.