acceptodds
Under review as a conference paper at ICLR 2027

Equal State Fit, Different Finite Responses in Vision-Language Models

Abstract

Cross-model steering uses maps between hidden states to reuse activation interventions in another model. We ask whether fitting paired states determines the effect of a mapped intervention on held-out target inputs. In an eight-dimensional residual subspace, four paired examples constrain at most six dimensions, leaving maps with the same calibration fit and target update norm that can produce different finite responses. We measure this difference by the response diameter, the range of target responses across these maps. Across three target vision-language models, 123 of 249 controlled groups have scalar diameters above a conservative numerical threshold. The separation recurs on image-disjoint Visual Question Answering v2 directions, including 19/24 Qwen-to-IDEFICS2 and 16/24 LLaVA-to-IDEFICS2 groups, and on a fixed Multimodal Visual Patterns subset spanning eight of nine visual patterns. All evaluable natural responses remain stable when the intervention step is halved, with further evidence from fixed angle grids and a second residual site. These results distinguish the role of state calibration from target-response evaluation: calibration fixes the map on the observed span, while updates outside that span require direct comparison at matched target norms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.