Activation Steering under Low-Bit Quantization: A Geometric Perspective on Degradation and Recovery
Abstract
Activation steering enables inference-time control of language-model behavior, but whether an existing steering interface survives low-bit quantization remains unclear. We study steering compatibility with the intervention layer, direction, and strength held fixed across full-precision and quantized models. We show that ordinary output fidelity alone does not guarantee preservation of responses to the same intervention. We explain the gap through the geometry of the quantization error. At the layer level, we relate directional response error to calibration geometry, showing that weakly covered directions admit larger worst-case errors under a fixed reconstruction budget. At the output level, we characterize intervention-dependent error amplification through the relative Fisher geometry of unsteered and intervened distributions. Under suitable coverage and bounded-sensitivity conditions, we derive a local upper bound on steered output KL in terms of unsteered output KL. Motivated by this analysis, we investigate post-quantization output distillation using only ordinary, unsteered teacher outputs. Experiments across multiple model families show that this procedure can improve steering compatibility without observing steering vectors or intervened outputs during training. These results identify steering controllability as a distinct deployment property of quantized language models and provide a geometric framework for analyzing compatibility failures and investigating their recovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.