acceptodds
Under review as a conference paper at ICLR 2027

Identifying, Installing, and Removing Capabilities in Vision-Language Models

Abstract

We study the computations installed by low-rank fine-tuning (LoRA) for translation, scale, and rotation tasks on a procedurally generated spatial benchmark in which one of two otherwise identical point clouds is transformed. Closed-form solutions to these simple stimuli provide both exact algorithms and cheaper statistics that are correct over part of the parameter range but fail over the rest. When asked in what direction one cloud of points must rotate to align with another, two 7B models fine-tuned on stimuli with rotation offsets between 0° and 45° do not calculate the answer based on correspondence between points, but instead compare the apparent orientation of each cloud. The fingerprint of this is evident on a full 0–180° set of evaluation stimuli, where performance alternates between high and inverted accuracy, with period 90° on statistically round clouds and 180° on elongated ones. A parameter-free reader matches these response curves by measuring the angle between the two clouds' principal axes and wrapping it into the symmetry period, agreeing with both model families. A full-range computation can be installed by training over the full 0–180° range at 1,200 examples per task when the points in the two clouds are labeled with matching numbers (cues), or at 18,000 examples per task on cue-free stimuli, where the model obtains one correspondence from the point farthest from the centroid. We also demonstrate that measured computations can be severed from model readout. A model competent at x (left/right), y (up/down), and their XOR can be tuned to emit a single constant token on x and y ("refuse" to answer), while XOR holds at perfect accuracy. The refused model's output still flips when either sign flips, so it remains sensitive to the quantities it will not verbalize. This severing costs a fraction of the supervision that installed the initial competencies. When framed as capability removal, our experiments demonstrate that, across multiple losses designed to provide readout suppression and/or algorithm destruction, such removals are successful but not durable. Reinstallation of removed verbalization or computation requires 13 to 800 times less supervision than initial installation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.