When Weight Fidelity Misleads: Concept-Erasure Safety Edits Under Diffusion-Model Quantization
Abstract
Concept-erasure edits are usually evaluated before quantization, although deployed models often run at lower numerical precision. We ask whether preservation of an edit in weight space predicts preservation of its function and behavior. On Stable Diffusion 1.5, we separately measure the deployed weight difference, the internal activation difference induced by the edit, and generated outputs. Under NF4 weight quantization, UCE's deployed weight difference has cosine similarity 0.568 with the original edit, while its activation-level effect remains aligned at 0.918. For the nudity edit, projecting the differential weight error onto edit-relevant conditioning inputs reduces its relative magnitude from 1.29 in weight space to 0.44. ESD-u also exhibits substantial distortion of the deployed weight difference while retaining high average activation alignment, extending the parameter–function decoupling to a different editing surface. We next quantize the text-conditioning inputs to selected cross-attention projections. At four activation bits, NudeNet efficacy falls by 0.264–0.290 for UCE and TIME, and CLIP measurements show the same direction, with all four paired intervals excluding zero. The INT8 weights remain identical across activation precisions. An exploratory 48-prompt, six-seed isolation with INT8 weights fixed attributes this change to activation rounding. Raw edited-image nudity scores also rise, matched-magnitude Gaussian controls do not reproduce the weakening, and Dreamlike Photoreal 2.0 reproduces the A4 effect with paired intervals excluding zero. On 50 unrelated prompts, CLIPScore declines similarly for base and edited models, providing no evidence of selective edited-model collapse under this probe. Our evidence is limited to one safety target, related SD-1.5 checkpoints, and simulated quantization of selected projections. Nevertheless, it shows that weight-space fidelity is insufficient: safety edits should be evaluated under the numerical configuration in which they will be deployed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.