Equal State Fit, Different Finite Responses in Vision-Language Models
Abstract
Cross-model steering uses maps between hidden states to reuse activation interventions in another model. We ask whether fitting paired states determines the effect of a mapped intervention on held-out target inputs. In an eight-dimensional residual subspace, four paired examples constrain at most six dimensions, leaving maps with the same calibration fit and target update norm that can produce different finite responses. We measure this difference by the response diameter, the range of target responses across these maps. Across three target vision-language models, 123 of 249 controlled groups have scalar diameters above a conservative numerical threshold. The separation recurs on image-disjoint Visual Question Answering v2 directions, including 19/24 Qwen-to-IDEFICS2 and 16/24 LLaVA-to-IDEFICS2 groups, and on a fixed Multimodal Visual Patterns subset spanning eight of nine visual patterns. All evaluable natural responses remain stable when the intervention step is halved, with further evidence from fixed angle grids and a second residual site. These results distinguish the role of state calibration from target-response evaluation: calibration fixes the map on the observed span, while updates outside that span require direct comparison at matched target norms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.