Stable Directions, Different Effects: Intervention Fidelity Across Model Precision
Abstract
Activation interventions can be characterized on one numerical realization of a language model and later applied to a quantized version. Preserving standard model quality or direction-level representation similarity does not guarantee that the intervention effect is preserved. We formalize this as cross-precision intervention fidelity and define a signed interventionfidelity gap between native and transferred steering effects. Across several language models, precision changes, model sizes, steering estimators, and prospectively held-out concepts/tasks, we compare a prespecified directionalignment/simple-geometry baseline with local functional sensitivity. Intervention effects can change even when directions remain strongly aligned, whereas a standard first-order local approximation tracks the signed gap closely in the strongest held-out evaluations. The novelty is not the Taylor expansion itself, but the intervention-fidelity problem and the empirical accuracy of this local account across the tested replication axes. A correction frozen before fresh evaluation improves reconstruction of native intervention effects and a prespecified sign-based decision. Finally, a prospective parsimony test favors the compact first-order approximation over the richer predictor. Within the tested normalized-transfer regime, intervention reliability should therefore be evaluated functionally rather than inferred from ordinary quantization quality or direction similarity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.