When Should a Steering Vector Transfer?Direction-Aware Transport with Selective Utility Certification
Abstract
Transferring activation directions from a large language model to a smaller model offers a way to modify behavior without updating either model’s weights. However, aligning representation spaces does not establish that a particular direction transfers, and uncertainty about an answer does not establish that an intervention will improve it. We formulate cross-model steering as selective deployment of a fixed candidate intervention. Our framework, VECTORGUARD, combines direction-specific transport diagnostics, a child-only predictor of paired intervention utility, and an independent calibration procedure that constrains both harmful changes and mean excess loss. The procedure can decline every intervention when the evidence is insufficient. Under independent and identically distributed calibration and deployment examples, we derive a finite-sample, population-conditional guarantee using simultaneous concentration bounds; it does not imply safety for every input or robustness to arbitrary distribution shift. We specify a staged evaluation separating baseline validity, transport quality, utility ranking, certification feasibility, and deployment cost. On GSM8K, selective deployment improves accuracy by 5.23 and 4.17 percentage points for the Qwen and cross-family pairs, respectively, compared with 0.61 and 0.23 points from unconditional intervention. The policies cover 43.6% and 40.6% of test inputs, with conditional harmful-flip rates below 1%, after certifying 45.0% and 40.5% calibration coverage. Mean latency rises by 11% and 10%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.